Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “large data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

University of Washington Participation in SAFARI-2000

This report presents a summary, of the participation in the 6-week field study in southern Africa. During the field study there were flown 119.02 research hours (31 research flights). In these flights the researchers obtained many unique data sets for evaluating the effects of biomass burning and other sources of particles and gases on the climate of southern Africa, and obtained simultaneous measurements with NASA ER-2 and Terra overflights. They also analyzed portions of the large data sets acquired in SAFARI-2000. They attended several SAFARI-2000 workshops, national and international conferences to present SAFARI-2000 results.

Hobbs, Peter V.↗

Seasonal to Decadal-Scale Variability in Satellite Ocean Color and Sea Surface Temperature for the California Current System

Support for this project was used to develop satellite ocean color and temperature indices (SOCTI) for the California Current System (CCS) using the historic record of CZCS West Coast Time Series (WCTS), OCTS, WiFS and AVHRR SST. The ocean color satellite data have been evaluated in relation to CalCOFI data sets for chlorophyll (CZCS) and ocean spectral reflectance and chlorophyll OCTS and SeaWiFS. New algorithms for the three missions have been implemented based on in-water algorithm data sets, or in the case of CZCS, by comparing retrieved pigments with ship-based observations. New algorithms for absorption coefficients, diffuse attenuation coefficients and primary production have also been evaluated. Satellite retrievals are being evaluated based on our large data set of pigments and optics from CalCOFI.

Mitchell, B. Greg↗

Methodology to Define Delivery Accuracy Under Current Day ATC Operations

In order to enable arrival management concepts and solutions in a NextGen environment, ground- based sequencing and scheduling functions have been developed to support metering operations in the National Airspace System. These sequencing and scheduling algorithms as well as tools are designed to aid air traffic controllers in developing an overall arrival strategy. The ground systems being developed will support the management of aircraft to their Scheduled Times of Arrival (STAs) at flow-constrained meter points. This paper presents a methodology for determining the undelayed delivery accuracy for current day air traffic control operations. This new method analyzes the undelayed delivery accuracy at meter points in order to understand changes of desired flow rates as well as enabling definition of metrics that will allow near-future ground automation tools to successfully achieve desired separation at the meter points. This enables aircraft to meet their STAs while performing high precision arrivals. The research presents a possible implementation that would allow delivery performance of current tools to be estimated and delivery accuracy requirements for future tools to be defined, which allows analysis of Estimated Time of Arrival (ETA) accuracy for Time-Based Flow Management (TBFM) and the FAA's Traffic Management Advisor (TMA). TMA is a deployed system that generates scheduled time-of-arrival constraints for en- route air traffic controllers in the US. This new method of automated analysis provides a repeatable evaluation of the delay metrics for current day traffic, new releases of TMA, implementation of different tools, and across different airspace environments. This method utilizes a wide set of data from the Operational TMA-TBFM Repository (OTTR) system, which processes raw data collected by the FAA from operational TMA systems at all ARTCCs in the nation. The OTTR system generates daily reports concerning ATC status, intent and actions. Due to its availability, ease of use, and vast collection of data across several airspaces it was determined that the OTTR data set would be the best method to utilize moving forward with this analysis. The particular variables needed for further analysis were determined along with the necessary OTTR reports, by working closely with the repository team additional analysis reports were developed that provided key ETA and STA information at the freeze horizon. One major benefit of the OTTR data is that using the correct reports the data across several airports could be analyzed over large periods of time. The OTTR data processes the TBFM data daily and is stored in various formats across several airspaces. This allowed us to develop our own parsing methods and raw data processing that would not rely on other computationally expensive tools that perform more in depth analysis of similar sets of data. The majority of this work consisted of the development of the ability to filter flights to create a subset of flights that could be considered undelayed, which is defined as a flight at the freeze horizon with an ETA and STA difference that was minimal or close to zero. This was a broad method that allowed the consideration of a large data set which consisted of all the traffic across a two month period in 2013, the hottest and coldest months, arriving into four airports: George Bush Intercontinental, Denver International, Los Angeles International, and Phoenix Sky Harbor.

delivery accuracy↗

Learning to Count Grave Sites for Cemetery Observation Models With Satellite Imagery

Understanding how people occupy open spaces is important for research in support of population modeling, policy, national security, emergency response, and sustainability. For the past decade, there has been an increase in research toward capturing and reporting population dynamics and patterns of life at the building level and in some open public spaces such as cemeteries and parks. This is done through observation models developed from local sociocultural information acquired at various spatiotemporal scales to inform night, day, and episodic population occupancy estimates (people/1000 sq ft). Sociocultural information for cemeteries and parks is scarcely available and often collected manually. The process is not only marred by inconsistencies but is laborious and time consuming. In this study, we leverage convolutional neural networks (CNNs) and satellite imagery to derive grave site counts as proxy variables to support scalable and accurate sociocultural data required in a population observation model. Through a hybrid workflow (weak localization plus regression model), we characterize a large scale automation process to counting of grave sites. We evaluate and demonstrate the efficacy of proposed workflow using out-of-data set large satellite imagery and establish its broader impact on cemetery observation models.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

The European Southern Observatory-MIDAS table file system

The new and substantially upgraded version of the Table File System in MIDAS is presented as a scientific database system. MIDAS applications for performing database operations on tables are discussed, for instance, the exchange of the data to and from the TFS, the selection of objects, the uncertainty joins across tables, and the graphical representation of data. This upgraded version of the TFS is a full implementation of the binary table extension of the FITS format; in addition, it also supports arrays of strings. Different storage strategies for optimal access of very large data sets are implemented and are addressed in detail. As a simple relational database, the TFS may be used for the management of personal data files. This opens the way to intelligent pipeline processing of large amounts of data. One of the key features of the Table File System is to provide also an extensive set of tools for the analysis of the final results of a reduction process. Column operations using standard and special mathematical functions as well as statistical distributions can be carried out; commands for linear regression and model fitting using nonlinear least square methods and user-defined functions are available. Finally, statistical tests of hypothesis and multivariate methods can also operate on tables.

Peron, M.↗

Getting to the Core of PARAFAC2, A Nonnegative Approach

In this paper, the authors present a novel method of performing PARAFAC2 factorization of three-way data using a compact representation of that data. In the standard PARAFAC2 algorithm, two modes of the data are recovered directly during the decomposition while the third mode is returned as a transformation matrix, which is then used to rotate sets of orthogonal third-mode basis factors into interpretable factors. In our new method, the data are first decomposed into a core matrix and orthogonal factor loading matrices in the first two modes as well as sets of orthogonal factors in the third mode (as in standard PARAFAC2). The core matrix is then decomposed using a the standard PARAFAC2 strategy to produce transformation matrices in all three modes. The algorithm is particularly useful for very large data sets and essentially permits imposition of nonnegativity in all three modes.

97 MATHEMATICS AND COMPUTING↗

Automated ISS Flight Utilities

During my internship at NASA Johnson Space Center, I worked in the Space Radiation Analysis Group (SRAG), where I was tasked with a number of projects focused on the automation of tasks and activities related to the operation of the International Space Station (ISS). As I worked on a number of projects, I have written short sections below to give a description for each, followed by more general remarks on the internship experience. My first project is titled "General Exposure Representation EVADOSE", also known as "GEnEVADOSE". This project involved the design and development of a C++/ ROOT framework focused on radiation exposure for extravehicular activity (EVA) planning for the ISS. The utility helps mission managers plan EVAs by displaying information on the cumulative radiation doses that crew will receive during an EVA as a function of the egress time and duration of the activity. SRAG uses a utility called EVADOSE, employing a model of the space radiation environment in low Earth orbit to predict these doses, as while outside the ISS the astronauts will have less shielding from charged particles such as electrons and protons. However, EVADOSE output is cumbersome to work with, and prior to GEnEVADOSE, querying data and producing graphs of ISS trajectories and cumulative doses versus egress time required manual work in Microsoft Excel. GEnEVADOSE automates all this work, reading in EVADOSE output file(s) along with a plaintext file input by the user providing input parameters. GEnEVADOSE will output a text file containing all the necessary dosimetry for each proposed EVA egress time, for each specified EVADOSE file. It also plots cumulative dose versus egress time and the ISS trajectory, and displays all of this information in an auto-generated presentation made in LaTeX. New features have also been added, such as best-case scenarios (egress times corresponding to the least dose), interpolated curves for trajectories, and the ability to query any time in the EVADES output. As mentioned above, GEnEVADOSE makes extensive use of ROOT version 6, the data analysis framework developed at the European Organization for Nuclear Research (CERN), and the code is written to the C++11 standard (as are the other projects). My second project is the Automated Mission Reference Exposure Utility (AMREU).Unlike GEnEVADOSE, AMREU is a combination of three frameworks written in both Python and C++, also making use of ROOT (and PyROOT). Run as a combination of daily and weekly cron jobs, these macros query the SRAG database system to determine the active ISS missions, and query minute-by-minute radiation dose information from ISS-TEPC (Tissue Equivalent Proportional Counter), one of the radiation detectors onboard the ISS. Using this information, AMREU creates a corrected data set of daily radiation doses, addressing situations where TEPC may be offline or locked up by correcting doses for days with less than 95% live time (the total amount time the instrument acquires data) by averaging the past 7 days. As not all errors may be automatically detectable, AMREU also allows for manual corrections, checking an updated plaintext file each time it runs. With the corrected data, AMREU generates cumulative dose plots for each mission, and uses a Python script to generate a flight note file (.docx format) containing these plots, as well as information sections to be filled in and modified by the space weather environment officers with information specific to the week. AMREU is set up to run without requiring any user input, and it automatically archives old flight notes and information files for missions that are no longer active. My other projects involve cleaning up a large data set from the Charged Particle Directional Spectrometer (CPDS), joining together many different data sets in order to clean up information in SRAG SQL databases, and developing other automated utilities for displaying information on active solar regions, that may be used by the space weather environment officers to monitor solar activity. I consulted my mentor Dr. Ryan Rios and Dr. Kerry Lee for project requirements and added features, and ROOT developer Edmond Offermann for advice on using the ROOT library. I also received advice and feedback from Dr. Janet Barzilla of SRAG, who tested my code. Besides these inputs, I worked independently, writing all of the code by myself. The code for all these projects is documented throughout, and I have attempted to write it in a modular format. Assuming that ROOT is updated accordingly, these codes are also Y2038-compliant (and Y10K-compliant). This allows the code to be easily referenced, modified and possibly repurposed for non-ISS missions in the future, should the necessary inputs exist. These projects have taught me a lot about coding and software design - I have become a much more skilled C++ programmer and ROOT user, and I also learned to code in Python and PyROOT (and its advantages and disadvantages compared to C++/ ROOT). Furthermore, I have learned about space radiation and radiation modeling, topics that greatly interest me as I pursue a degree in physics. Working alongside experimental physicists like Dr. Rios, I have developed a greater understanding and appreciation for experimental science, something I have always leaned towards but to which I lacked significant exposure. My work in SRAG has also given me the invaluable opportunity to witness the work environment for physicists at NASA, and what a career in academia may look like at a government laboratory such as NASA Johnson Space Center. As I continue my studies and look forward to graduate school and a future career, this experience at NASA has given me a meaningful and enjoyable opportunity to put my skills to use and see what my future career path might hold.

Offermann, Jan Tuzlic↗

Data simulation for the Lightning Imaging Sensor (LIS)

This project aims to build a data analysis system that will utilize existing video tape scenes of lightning as viewed from space. The resultant data will be used for the design and development of the Lightning Imaging Sensor (LIS) software and algorithm analysis. The desire for statistically significant metrics implies that a large data set needs to be analyzed. Before 1990 the quality and quantity of video was insufficient to build a usable data set. At this point in time, there is usable data from missions STS-34, STS-32, STS-31, STS-41, STS-37, and STS-39. During the summer of 1990, a manual analysis system was developed to demonstrate that the video analysis is feasible and to identify techniques to deduce information that was not directly available. Because the closed circuit television system used on the space shuttle was intended for documentary TV, the current value of the camera focal length and pointing orientation, which are needed for photoanalysis, are not included in the system data. A large effort was needed to discover ancillary data sources as well as develop indirect methods to estimate the necessary parameters. Any data system coping with full motion video faces an enormous bottleneck produced by the large data production rate and the need to move and store the digitized images. The manual system bypassed the video digitizing bottleneck by using a genlock to superimpose pixel coordinates on full motion video. Because the data set had to be obtained point by point by a human operating a computer mouse, the data output rate was small. The loan and subsequent acquisition of a Abekas digital frame store with a real time digitizer moved the bottleneck from data acquisition to a problem of data transfer and storage. The semi-automated analysis procedure was developed using existing equipment and is described. A fully automated system is described in the hope that the components may come on the market at reasonable prices in the next few years.

Boeck, William L.↗

A geospatial risk analysis graphical user interface for identifying hazardous chemical emission sources

Background: Performing back trajectory and forward trajectory using the Hybrid Single-Particle Lagrangian Integrated Trajectory Model (HYSPLIT) is a reliable approach for assessing particle transport after release among mid-field atmospheric models. HYSPLIT has an externally facing online interface that allows non-expert users to run the model trajectories without requiring extensive training or programming. However, the existing HYSPLIT interface is limited if simulations have a large amount of meteorological data and timesteps that are not coincident. The objective of this study is to design and develop a more robust tool to rapidly evaluate hazard transport conditions and to perform risk analysis, while still maintaining an intuitive and user-friendly interface. Methods: HYSPLIT calculates forward and backward trajectories of particles based on wind speed, wind direction, and the corresponding location, timestamp, and Pasquill stability classes of the regions of the atmosphere in terms of the wind speed, the amount of solar radiation, and the fractional cloud cover. The computed particle transport trajectories, combined with the online Proton Transfer Reaction-Mass Spectrometry (PTR-MS) data (https://figshare.com/articles/dataset/ARL_Data_from_PROS_station_at_Hanford_site/19993964), can be used to identify and quantify the sources and affected area of the hazardous chemicals’ emission using the potential source distribution function (PSDF). PSDF is an improved statistical function based on the well-known potential source contribution function (PSCF) in establishing the air pollutant source and receptor relationship. Performing this analysis requires a range of meteorological and pollutant concentration measurements to be statistically meaningful. The existing HYSPLIT graphical user interface (GUI) does not easily permit computations of trajectories of a dataset of meteorological data in high temporal frequency. To improve the performance of HYSPLIT computations from a large dataset and enhance risk analysis of the accidental release of material at risk, a geospatial risk analysis tool (GRAT-GUI) is created to allow large data sets to be processed instantaneously and to provide ease of visualization. Results: The GRAT-GUI is a native desktop-based application and can be run in any Windows 10 system without any internet access requirements, thus providing a secure way to process large meteorological datasets even on a standalone computer. GRAT-GUI has features to import, integrate, and convert meteorological data with various formats for hazardous chemical emission source identification and risk analysis as a self-explanatory user interface. The tool is available at https://figshare.com/articles/software/GRAT/19426742.

97 MATHEMATICS AND COMPUTING↗

A Multi-Band Analytical Algorithm for Deriving Absorption and Backscattering Coefficients from Remote-Sensing Reflectance of Optically Deep Waters

A multi-band analytical (MBA) algorithm is developed to retrieve absorption and backscattering coefficients for optically deep waters, which can be applied to data from past and current satellite sensors, as well as data from hyperspectral sensors. This MBA algorithm applies a remote-sensing reflectance model derived from the Radiative Transfer Equation, and values of absorption and backscattering coefficients are analytically calculated from values of remote-sensing reflectance. There are only limited empirical relationships involved in the algorithm, which implies that this MBA algorithm could be applied to a wide dynamic range of waters. Applying the algorithm to a simulated non-"Case 1" data set, which has no relation to the development of the algorithm, the percentage error for the total absorption coefficient at 440 nm a (sub 440) is approximately 12% for a range of 0.012 - 2.1 per meter (approximately 6% for a (sub 440) less than approximately 0.3 per meter), while a traditional band-ratio approach returns a percentage error of approximately 30%. Applying it to a field data set ranging from 0.025 to 2.0 per meter, the result for a (sub 440) is very close to that using a full spectrum optimization technique (9.6% difference). Compared to the optimization approach, the MBA algorithm cuts the computation time dramatically with only a small sacrifice in accuracy, making it suitable for processing large data sets such as satellite images. Significant improvements over empirical algorithms have also been achieved in retrieving the optical properties of optically deep waters.

Lee, Zhong-Ping↗

A new Monte Carlo generator for BSM physics in B → K*ℓ+ℓ− decays with an application to lepton non-universality in angular distributions

Abstract Within the widely used EvtGen framework, we have added a new event generator model forB → K * ℓ + ℓ − with improved standard model (SM) decay amplitudes and possible BSM physics contributions, which are implemented in the operator product expansion in terms of Wilson coefficients. This event generator can then be used to estimate the statistical sensitivity of a simulated experiment to the most general BSM signal resulting from dimension-six operators. We describe the advantages and potential of the newly developed ‘Sibidanov Physics Generator’ in improving the experimental sensitivity of searches for lepton non-universal BSM physics and clarifying signatures. The new generator can properly simulate BSM scenarios, interference between SM and BSM amplitudes, and correlations between different BSM observables as well as acceptance bias. We show that exploiting such correlations substantially improves experimental sensitivity. As a demonstration of the utility of the MC generator, we examine the prospects for improved measurements of lepton non-universality in angular distributions forB→K * ℓ + ℓ − decays from the expected 50 ab −1 data set of the Belle II experiment, using a four-dimensional unbinned maximum likelihood fit. We describe promising experimental signatures and correlations between observables. The use of lepton-universality violating ∆-observables significantly reduces uncertainties in the SM expectations due to QCD and resonance effects and is ideally suited for Belle II with the large data sets expected in the next decade. Thanks to the clean experimental environment of ane + e − machine, Belle II should be able to probe BSM physics in the Wilson coefficientsC 7 and$$ {C}_7^{\prime } $$ C 7 ′ , which appear at lowq 2 in the di-electron channel.

Physics↗

ParChain: a framework for parallel hierarchical agglomerative clustering using nearest-neighbor chain

This paper studies the hierarchical clustering problem, where the goal is to produce a dendrogram that represents clusters at varying scales of a data set. We propose the ParChain framework for designing parallel hierarchical agglomerative clustering (HAC) algorithms, and using the framework we obtain novel parallel algorithms for the complete linkage, average linkage, and Ward's linkage criteria. Compared to most previous parallel HAC algorithms, which require quadratic memory, our new algorithms require only linear memory, and are scalable to large data sets. ParChain is based on our parallelization of the nearest-neighbor chain algorithm, and enables multiple clusters to be merged on every round. We introduce two key optimizations that are critical for efficiency: a range query optimization that reduces the number of distance computations required when finding nearest neighbors of clusters, and a caching optimization that stores a subset of previously computed distances, which are likely to be reused. Experimentally, we show that our highly-optimized implementations using 48 cores with two-way hyper-threading achieve 5.8--110.1x speedup over state-of-the-art parallel HAC algorithms and achieve 13.75--54.23x self-relative speedup. Compared to state-of-the-art algorithms, our algorithms require up to 237.3x less space. Our algorithms are able to scale to data set sizes with tens of millions of points, which existing algorithms are not able to handle.

Computer Science↗

Implementing Access to Data Distributed on Many Processors

A reference architecture is defined for an object-oriented implementation of domains, arrays, and distributions written in the programming language Chapel. This technology primarily addresses domains that contain arrays that have regular index sets with the low-level implementation details being beyond the scope of this discussion. What is defined is a complete set of object-oriented operators that allows one to perform data distributions for domain arrays involving regular arithmetic index sets. What is unique is that these operators allow for the arbitrary regions of the arrays to be fragmented and distributed across multiple processors with a single point of access giving the programmer the illusion that all the elements are collocated on a single processor. Today's massively parallel High Productivity Computing Systems (HPCS) are characterized by a modular structure, with a large number of processing and memory units connected by a high-speed network. Locality of access as well as load balancing are primary concerns in these systems that are typically used for high-performance scientific computation. Data distributions address these issues by providing a range of methods for spreading large data sets across the components of a system. Over the past two decades, many languages, systems, tools, and libraries have been developed for the support of distributions. Since the performance of data parallel applications is directly influenced by the distribution strategy, users often resort to low-level programming models that allow fine-tuning of the distribution aspects affecting performance, but, at the same time, are tedious and error-prone. This technology presents a reusable design of a data-distribution framework for data parallel high-performance applications. Distributions are a means to express locality in systems composed of large numbers of processor and memory components connected by a network. Since distributions have a great effect on the performance of applications, it is important that the distribution strategy is flexible, so its behavior can change depending on the needs of the application. At the same time, high productivity concerns require that the user be shielded from error-prone, tedious details such as communication and synchronization.

James, Mark↗

Satellite/rocket ozone comparisons at Natal, Brazil

Comparisons are presented of satellite, rocket, and balloon ozone profiles near Natal, Brazil (5.9 deg S, 35.2 deg W). The low variability of stratospheric ozone at Natal during March and April of 1985 has allowed intercomparisons of reasonably large data sets, rather than a small number of paired satellite/in situ comparisons. There are sharp differences between the profile from the SBUV instrument on Nimbus 7 and the in situ measurements. These results support the conclusions of the NASA Ozone Trends Panel that there is an instrumental cause for the very large changes in upper stratospheric ozone seen by SBUV. Along with other comparisons, these results are being used in a reassessment of the SBUV instrument and its data reduction procedures. The agreement between the ozone profiles from the SAGE II instrument on the ERBS satellite and the rocket values is excellent over the full range of comparisons. Both SAGE II and ROCOZ-A must convert from altitude to pressure for intercomparisons with SME and with SBUV-type instruments. The conversion between pressure and altitude is as important as the ozone measurements, especially in the upper stratosphere where the scale height for ozone is approximately half that for pressure.

Barnes, Robert A.↗

The scientific potential and technological challenges of the High-Luminosity Large Hadron Collider program

Here, we present an overview of the High-Luminosity (HL-LHC) program at the Large Hadron Collider (LHC), its scientific potential and technological challenges for both the accelerator and detectors. The HL-LHC program is expected to start circa 2027 and aims to increase the integrated luminosity delivered by the LHC by an order of magnitude at the collision energy of 14 TeV. This requires upgrades to the injector system, accelerator complex and luminosity levelling. The two experiments, ATLAS and CMS, require substantial upgrades to most of their systems in order to cope with the increased interaction rate, and much higher radiation levels than at the current LHC. We present selected examples based on novel ideas and technologies for applications at a hadron collider. Both experiments will replace their tracking systems. We describe the ATLAS pixel detector upgrade featuring novel tilted modules, and the CMS Outer Tracker upgrade with a new module design enabling use of tracks in the level-1 trigger system. CMS will also install state-of-the-art highly segmented calorimeter endcaps. Finally, we describe new picosecond precision timing detectors of both experiments. In addition, we discuss how the upgrades will enhance the physics performance of the experiments, and solve the computing challenges posed by the expected large data sets. The physics program of the HL-LHC is focused on precision measurements probing the limits of the Standard Model (SM) of particle physics and discovering new physics. We present a selection of studies that have been carried out to motivate the HL-LHC program. A central topic of exploration will be the characterization of the Higgs boson. The large HL-LHC data samples will extend the sensitivity of searches for new particles or new interactions whose existence has been hypothesized in order to explain shortcomings of the SM. Finally, we comment on the nature of large scientific collaborations.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

NASA SERC 1990 Symposium on VLSI Design

This document contains papers presented at the first annual NASA Symposium on VLSI Design. NASA's involvement in this event demonstrates a need for research and development in high performance computing. High performance computing addresses problems faced by the scientific and industrial communities. High performance computing is needed in: (1) real-time manipulation of large data sets; (2) advanced systems control of spacecraft; (3) digital data transmission, error correction, and image compression; and (4) expert system control of spacecraft. Clearly, a valuable technology in meeting these needs is Very Large Scale Integration (VLSI). This conference addresses the following issues in VLSI design: (1) system architectures; (2) electronics; (3) algorithms; and (4) CAD tools.

Maki, Gary K.↗

Machine Learning Approaches to Predicting Induced Seismicity and Imaging Geothermal Reservoir Properties

This project developed machine learning (ML) methods, lab data sets, and field data to advance geothermal exploration and geothermal energy production. The work had three focus areas. One involved the development of ML methods to use microearthquakes (MEQs) for imaging geothermal reservoir properties and improving subsurface characterization – most importantly the evolution of permeability within the evolving reservoir. This part of the work included development of ML approaches for automated MEQ location, focal mechanism determination and identification of earthquake precursors. The second area focused on using MEQ signals generated by geothermal exploration and production to predict the relationship between fluid injection and seismicity. Here, we extended to reservoir scale our success in using ML to predict laboratory earthquakes and fault zone stress state. The third focus area was on lab experiments. Here, we developed new ML models for lab earthquake prediction and identification of precursors to failure to improve earthquake forecasting and early warning in geothermal settings. Major outcomes of our work include ML models that learn from MEQ signals during geothermal exploration and production to predict induced seismicity. MEQs occur naturally in connection with drilling and energy production. We developed ML methods to use the seismic waves from these events to characterize the elastic, hydraulic and poromechanical properties of reservoirs. Our work illuminated fracture geometry and the evolution of fracture permeability by incorporating seismic coda wave analysis and ML methods to relate fluid injection and seismicity. We significantly expanded laboratory earthquake prediction to include methods that use both passive measurements of microearthquakes within the lab fault zones and also active source acoustic measurements of fault zone elastic properties. These methods can now predict fault zone stress state, time to failure and the magnitude of lab earthquakes. Our work showed that repetitive stick- slip failure events during frictional sliding (the lab equivalent of earthquakes) are preceded by a cascade of micro-failure events that radiate energy in a manner that foretells unstable failure – manifest as laboratory MEQs. We documented a mapping between fracture properties and statistical attributes of elastic radiation. We extended existing works to geothermal reservoir scale and developed ML methods to determine reservoir permeability, fracture properties, and their evolution during geothermal energy production. An attractive feature of ML algorithms is their ability to handle big datasets and reveal patterns and correlations that may remain invisible to conventional analyses. Our work connected data from field, laboratory and intermediate scales to study permeability, stress, strength, fracture stiffness and geometry. At the field scale we used data from the Newberry Volcano field site, UtahFORGE, EGS Collab, and also the Bedretto underground research lab in Switzerland. These data sets are bridging the gap between the lab scale, theory, and reservoir scale. Our work produced plain language summaries to improve public understanding of DOE research. We also developed openly distributed ML and seismicity datasets for use by all researchers and we published connections between induced seismicity in geothermal areas and reservoir properties including permeability, fracture properties, and stress state. Our models are designed for the large data sets of induced seismicity typically associated with geothermal sites. We produced labeled event catalogs and used them on geothermal data to assess how ML can facilitate geothermal production and exploration. All datasets are available on the GDR Productivity: The project produced 32 publications in peer reviewed journals (two are in review). It supported the work of 6 PhD students, 40 conference presentations, 6 keynote talks at national meetings, and mentoring and professional development for 4 postdoctoral fellows.

15 GEOTHERMAL ENERGY↗

Large Field Visualization with Demand-Driven Calculation

We present a system designed for the interactive definition and visualization of fields derived from large data sets: the Demand-Driven Visualizer (DDV). The system allows the user to write arbitrary expressions to define new fields, and then apply a variety of visualization techniques to the result. Expressions can include differential operators and numerous other built-in functions, ail of which are evaluated at specific field locations completely on demand. The payoff of following a demand-driven design philosophy throughout becomes particularly evident when working with large time-series data, where the costs of eager evaluation alternatives can be prohibitive.

Moran, Patrick J.↗