Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data exploration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Carbon Storage Site Mapping Inquiry Tool (MapIT)

The Carbon Storage Site Mapping Inquiry Tool (MapIT) is an online web mapping tool designed to help users discover available public-sourced data to facilitate data exploration in support of Underground Injection Control (UIC) Program Class VI Well Site permitting for the geologic sequestration of carbon dioxide.

Pantaleone, Scott↗

Identification of Preferential Recharge Zones in Karst Systems Based on the Correlation between the Spring Level and Precipitation: A Case Study from Jinan Spring Basin

The Jinan spring basin is located in the karst area of northern China, where springs serve as important sources of water supply. Several studies on spring protection and water supply have been carried out, and scholars have developed some laws on local groundwater flow dynamic and characteristics of aquifer structures. Unfortunately, there is a lack of detailed research on preferential recharge zones, which are the main recharge pathways of springs. Therefore, this research focuses on identifying preferential recharge zones based on the correlation between the spring level and precipitation. The results show that when precipitation is more intense or lasts longer, there is a stronger correlation between spring level and precipitation. It has been established that the precipitation at Donghongmiao station has the closest relationship with the dynamic of Baotu spring, which is found to be the most significant contribution to spring preservation. Two potential preferential recharge zones in the Jinan spring basin are detected through correlation analysis and geological exploration data. These findings support spring protection and water supply projects in karst regions.

58 GEOSCIENCES↗

Spectral Properties of Globally Distributed ENA Fluxes across Diverse Regions of the Heliosphere

This study analyzes energetic neutral atom (ENA) spectral properties across distinct regions of globally distributed flux (GDF) sky maps, using Interstellar Boundary Explorer data from a full solar cycle, corrected for time dispersion. By time-shifting the data to the heliosheath using GDF source distances from D. B. Reisenfeld et al., we achieve a more accurate representation of heliosheath GDF energy spectra. We quantify ENA spectral characteristics, heliosheath line-of-sight-integrated proton pressure, and heliosheath proton temperature, comparing these to solar wind properties at 1 au and interplanetary scintillation-derived solar wind data. Our findings show that the spectral index is generally anticorrelated with heliosheath proton temperature and pressure, except in the central tail, where a partial positive correlation is observed. The lowest spectral index values occur when high-latitude heliosheath regions are dominated by fast solar wind from polar coronal holes. The south pole exhibits the flattest energy spectra due to plasma heating from both fast solar wind and a late-2014 pressure pulse. The central tail shows shorter variability (5–6 yr) for spectral index and heliosheath proton temperature, while proton pressure follows the 11 yr solar cycle. Most spectral shapes exhibit a “knee” distribution, peaking during solar maximum, with an “ankle” shape observed only at the south pole during solar cycle transitions. Asymmetry in proton pressure in the lobes is driven by the draping effect of the local interstellar magnetic field. This study provides insights into the energetic properties of GDF across the heliosphere, enhancing our understanding of the heliospheric environment.

79 ASTRONOMY AND ASTROPHYSICS↗

DELVE Milky Way Satellite Galaxy Census. I. Satellite Population and Survey Selection Function in DES, DELVE, and Pan-STARRS

The properties of Milky Way satellite galaxies have important implications for galaxy formation, reionization, and the fundamental physics of dark matter. However, the population of Milky Way satellites includes the faintest known galaxies, and current observations are incomplete. To understand the impact of observational selection effects on the known satellite population, we perform rigorous, quantitative estimates of the Milky Way satellite galaxy detection efficiency in three wide-field survey datasets: the Dark Energy Survey Year 6, the DECam Local Volume Exploration Data Release 3, and the Pan-STARRS1 Data Release 1. Together, these surveys cover ∼13,600 deg 2 to g ∼ 24.0 and ∼27,700 deg 2 to g ∼ 22.5, spanning ∼91% of the high-Galactic-latitude sky (∣b∣ ≥ 15°). We apply multiple detection algorithms over the combined footprint and recover 49 known satellites above a strict census detection threshold. To characterize the sensitivity of our census, we run our detection algorithms on a large set of simulated galaxies injected into the survey data, which allows us to develop models that predict the detectability of satellites as a function of their properties. We then fit an empirical model to our data and infer the luminosity function, radial distribution, and size–luminosity relation of Milky Way satellite galaxies. Our empirical model predicts a total of $265^{+79}_{-47}$ satellite galaxies with −20 ≤ M V ≤ 0, half-light radii of 15 ≤ r 1/2 , (pc) ≤ 3000, and galactocentric distances of 10 ≤ D GC (kpc) ≤ 300. We also identify a mild anisotropy in the angular distribution of the observed galaxies, at a significance of ∼2σ, which can be attributed to the clustering of satellites associated with the LMC.

Tan, Chin Yi [Univ. of Chicago, IL (United States)↗

Nuclear data for space exploration

Understanding the harmful effects of galactic cosmic rays (GCRs) on space exploration requires a substantial amount of nuclear data. Specifically, the interaction of energetic GCR charged particles with spacecraft materials generates secondary radiations that, through energy deposition, can harm astronauts and electronic systems. By identifying the gaps in our knowledge of the relevant nuclear data—such as interaction cross sections—and identifying ways to fill those gaps—with measurements, compilations, evaluations, dissemination, reaction modeling, sensitivity studies, and uncertainty quantification—the safety and viability of space exploration can be improved. This work surveys the state of the art in this interdisciplinary field and identifies promising collaborative research topics that have significant potential to advance our understanding of the effects of the space radiation environment on space exploration.

42 ENGINEERING↗

Exploring Wildfire & Energy data toward State Prioritization Index (WESPI)

Energy infrastructure can both induce and suffer risks from wildfires ranging from direct damage to energy assets such as substations and power lines to Public Safety Power Shutoffs. Recent wildfire events underscore the need for data-driven approaches that help states and utilities proactively plan for wildfire risk. Existing national tools such as Federal Emergency Management Agency (FEMA)’s National Risk Index (NRI) are valuable for community hazard planning. However, they are less suited for energy infrastructure, as they emphasize population and building exposure rather than system vulnerabilities. In this paper, we explore relationships between energy and wildfire data and present a Wildfire-Energy State Prioritization Index (WESPI). Our methodology combines data from the US Forest Service’s Fire Simulation (FSIM) dataset with energy resilience metrics, historical fire incidents, and geospatial data on transmission lines and fire stations. Correlation analyses suggest that FSIM burn probability is more strongly associated with power outage metrics (ρ = 0.32) than NRI wildfire frequency, and counties with a greater density of fire stations experience more frequent, but less intense wildfires. We further leverage data for burn probability, transmission line density, and fire station density to develop a Wildfire-Energy State Prioritization Index (WESPI) to highlight counties where wildfire hazard, infrastructure exposure, and limited suppression capacity converge. The index provides a consistent, scalable framework for state energy offices and utilities to screen counties for vegetation management, optimization of outage management system deployment, and to inform wildfire mitigation plans.

Critical infrastructure↗

Exploring Continuous Seismic Data at an Industry Facility Using Unsupervised Machine Learning

Seismic data recorded at industrial sites contain valuable information on anthropogenic activities. With advances in machine learning and computing power, new opportunities have emerged to explore the seismic wavefield in these complex environments. We applied two unsupervised machine learning algorithms to analyze continuous seismic data collected from an industrial facility in Texas, United States. The Uniform Manifold Approximation and Projection for Dimension Reduction algorithm was used to reduce the dimensionality of the data and generate 2D embeddings. Then, the Hierarchical Density-Based Spatial Clustering of Applications with Noise method was employed to automatically group these embeddings into distinct signal clusters. Our analysis of over 1400 hr (around 59 days) of continuous seismic data revealed five and seven signal clusters at two separate stations. At both stations, we identified clusters associated with background noise and vehicle traffic, with the latter’s temporal patterns aligning closely with the facility’s work schedule. Furthermore, the algorithms detected signal clusters from unknown sources and underline the ability of unsupervised machine learning for uncovering previously unrecognized patterns. Our analysis demonstrates the effectiveness of unsupervised approaches in examining continuous seismic data without requiring prior knowledge or pre-existing labels.

58 GEOSCIENCES↗

AI-Batt (Autonomous Identification of Battery Life Models) [SWR 21-36]

Autonomous Identification of Battery Life Models (AI-Batt) AI-Batt is a MATLAB code base for developing lifetime models for batteries from accelerated aging data. The code base provides many functions for processing, visualizing, and modeling battery aging data, making the data processing, exploration, and modeling workflow substantially faster. These tools are tailored for working with battery aging data sets, which usually consist of many separate time-series for each cell, with many test conditions and possible replicates at each condition, which makes it difficult to simply process or visualize the data set. Complex modeling tasks, such as cross-validation, sensitivity analysis, and uncertainty quantification have been implemented to enable thorough statistical investigation of model predictions. Additionally, several machine-learning algorithms are implemented to autonomously identify suitable models via symbolic regression. Data processing functions automatically cast data from the struct data type, which is commonly used to store experimental data, but is not an acceptable input for most algorithms, to the table data type, which can be easily used as input to any optimization algorithm. Also, the data can be separated into time-invariant and time-variant data tables, which is helpful for exploring the data set as well as developing separate models for time-variant and time-invariant aging mechanisms. For example, in aging tests with constant temperature, temperature is a time-invariant experimental condition. Visualization tools enable plotting of data, model fits, and model simulations possible with single-line function calls, empowering data exploration of complex data sets with both time-varying and time-invariant trends. Plots can be automatically generated for the whole data set, or separated by data group (groups of test replicates) or individual data series. Data points or data series can be automatically colored by the value of a variable with a variety of color maps, and model predictions can also be colored by the value of a fit statistic. Comparisons between data sets and the predictions/simulations of different models on the same data set can be easily plotted as well. Distributions of parameter values from bootstrap resampling can be plotted to visualize the reliability of parameter estimation, or determine any correlations between parameters. Modeling tools handle the complex task of creating and parsing symbolic equations for modeling battery lifetime. Equations are parsed to grab relevant data variables, parameter values, or specified sub-models for input into optimization, evaluation, or simulation functions. Models can be optimized locally (one set of parameters for each data series), bi-level (some parameters shared across the data set), or globally (single set of parameters for all data). Functions implementing symbolic regression algorithms help users to discover effective model equations, even in poorly sampled, high-dimensional data.

Smith, Kandler↗

Divide and conquer: using RhizoVision Explorer to aggregate data from multiple root scans using image concatenation and statistical methods

Roots are important in agricultural and natural systems for determining plant productivity and soil carbon inputs. Sometimes, the amount of roots in a sample is too much to fit into a single scanned image, so the sample is divided among several scans, and there is no standard method to aggregate the data. Here, we describe and validate two methods for standardizing measurements across multiple scans: image concatenation and statistical aggregation. We developed a Python script that identifies which images belong to the same sample and returns a single, larger concatenated image. These concatenated images and the original images were processed with RhizoVision Explorer, a free and open-source software. An R script was developed, which identifies rows of data belonging to the same sample and applies correct statistical methods to return a single data row for each sample. These two methods were compared using example images from switchgrass, poplar, and various tree and ericaceous shrub species from a northern peatland and the Arctic. Most root measurements were nearly identical between the two methods except median diameter, which cannot be accurately computed by statistical aggregation. We believe the availability of these methods will be useful to the root biology community.

59 BASIC BIOLOGICAL SCIENCES↗

Big Data Meets Geothermal Exploration (CRADA Final Report)

As part of the Cyclotron Road program, Zanskar Geothermal & Minerals, Inc. investigated the application of micro-earthquake and ambient noise seismology methods to imaging and characterizing the structural characteristics and hydrothermal flux of subsurface faults. Significant advances in what could be resolved were enabled by two major developments in seismology: 1) the availability of large-n arrays of low-cost seismometers, and 2) the availability of increased computational power and semi-automated data reduction algorithms. In tandem, these advances may improve the signal-to-noise ratio and spatial precision of the data collected and enable higher-resolution characterization of subsurface fracture systems and their spatio-temporal evolution. These tools supported efforts to reduce dry-hole risk and to improve wellfield productivity for geothermal resource development. In particular, two applications of these advances were evaluated: 1) fracture-seismic imaging, which was used to detect ambient emissions from fluid-filled fractures, and 2) reservoir tomography, which used information about travel paths, source locations, and source parameters of micro-earthquakes to identify areas of enhanced permeability. Integration of these methods provided guidance for siting wells and served as prior constraints for reservoir models, informing forecasts of power potential and production and injection strategies aimed at minimizing temperature decline and improving overall resource productivity.

15 GEOTHERMAL ENERGY↗

Scalable and Energy-Efficient Methods for Interactive Exploration of Scientific Data

The main scientific contributions of this project are the following novel concepts for multidimensional arrays: shape-based similarity join (SIGMOD 2016), incremental view maintenance (SIGMOD 2017), user-defined stencil functions (HPDC 2017), and distributed caching for in-situ processing (SSDBM 2018). Building on our collaboration with the astrophysics group at LBNL, we applied these techniques to the data generated in the Palomar Transient Factory (PTF) astronomical survey. They played a pivotal role in the first-ever observation of a neutron star merger, which produces gravitational waves and turns out to be the origin of heavy elements, including gold. This has lead to a Science magazine article that has received extensive media coverage on ACM TechNews, Slashdot, FiveThirtyEight, and Quanta Magazine, among others. Additionally, two other articles detailing related aspects of the same discovery have been published in the Astrophysical Journal Letters journal. These publications have more than 3,000 citations according to Google Scholar (as of February 2022). This cross-disciplinary collaboration provided very good opportunities to apply database techniques to real-life scientific problems. The fact that they facilitated major discoveries in astrophysics proves the importance of our research. In addition to the work on multidimensional array databases, this project has also developed stochastic gradient descent (SGD) optimization algorithms for training large scale machine learning models, methods for querying in-situ data, and a database query optimizer based on sketch synopses.

79 ASTRONOMY AND ASTROPHYSICS↗

Exploring MDSplus data-acquisition software and custom devices

MDSplus is a software tool designed for data acquisition, storage, and analysis of complex scientific experiments. Over the years, MDSplus has primarily been used for data management for fusion experiments. This paper demonstrates that MDSplus can be used for a much wider variety of systems and experiments. We present a step-by-step tutorial describing how to create a simple experiment, manage the data, and analyze it using MDSplus and Python. To this end, a custom example device was developed to be used as the data source. This device was built on an opensource electronic hardware platform, and it consists of a microcontroller and two sensors. We read data from these sensors, store it in MDSplus, and use JupyterLab to visualize and process it. This project and code demo are available on the GitHub site at this URL: https://github.com/santorofer/MDSplusAndCustomeDevices

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Expansion-history preferences of DESI DR2 and external data

We explore the origin of the preference of Dark Energy Spectroscopic Instrument (DESI) Data Release 2 (DR2) baryon acoustic oscillation measurements and external data from cosmic microwave background (CMB) and type Ia supernovae (SNIa) that dark energy behavior departs from that expected in the standard cosmological model with vacuum energy (Λ ⁢CDM). In our analysis, we allow a flexible scaling of the expansion rate with redshift that nevertheless allows reasonably tight constraints on the quantities of interest, and adopt and validate a simple yet accurate compression of the CMB data that allows us to constrain our phenomenological model of the expansion history. We find that data consistently show a preference for a 3%–4% increase in the expansion rate at 𝑧 ≃ 0.7 relative to that predicted by the standard Λ⁢ CDM model, in excellent agreement with results from the less flexible (𝑤 0 ,𝑤 𝑎 ) parametrization which was used in previous analyses. Even though our model allows a departure from the best-fit Λ⁢ CDM model at zero redshift, we find no evidence for such a signal. We also find no evidence (at greater than 1⁢𝜎 significance) for a departure of the expansion rate from the Λ ⁢CDM predictions at higher redshifts for any of the data combinations that we consider. Altogether, our results strengthen the robustness of the findings using the combination of DESI, CMB, and SNIa data to dark-energy modeling assumptions.

Cosmological parameters↗

Performance Debugging and Tuning of Flash-X with Data Analysis Tools

State-of-the-art multiphysics simulations running on large scale leadership computing platforms have many variables contributing to their performance and scaling behavior. We recently encountered an interesting performance anomaly in Flash-X, a multiphysics multicomponent simulation software, when characterizing its performance behavior on several large-scale HPC platforms. The anomaly was tracked down to the interaction between the use of dynamic allocation of scratch data and data locality in the cache hierarchy. In this paper we present the details of unexpected performance variability of Flash-X, its extensive analysis using the performance measurement tool TAU to collect the data and Python data analysis libraries to explore the data, and our insights from this experience. In this process, we discovered and removed or mitigated two additional performance limiting bottlenecks for performance tuning.

Huck, Kevin↗

University Data Management Pilot Utilizing the Nuclear Research Data System

Background In 2022, the Office of Science and Technology Policy (OSTP) issued a memo that significantly reshaped the landscape of access to federally funded research. The memo mandated that all taxpayer-funded research be made available to the public without delay upon publication, without an embargo period, superseding the 2013 OSTP public access policy. This public access policy promotes transparency and the democratization of knowledge, ensuring that the fruits of scientific endeavors funded by federal agencies could be immediately accessed and built upon by scientists, educators, students, and the public at large. To implement the requirements of the OSTP guidance and DOE Public Access Plan, the Office of Nuclear Energy (NE) has implemented public access plan guidance and has identified several areas where better data management practices would further expand public access to important nuclear energy related scientific data, reports, and other technical products. Significant NE supported efforts are already underway for data management and public access to important nuclear energy related data.1 2 To address gaps in data management practices, and improve retention and accessibility of data, NE is actively exploring enhanced data management options utilizing its high-performance computing resources administered by its Nuclear Scientific User Facility Program. A newly piloted system, the Nuclear Research Data System (NRDS) acts as a portal for data collection and dissemination. Nuclear Energy University Program Research and Development Portfolio According to Web of Science, NEUP has produced 2,345 journal publication that have been cited more than 61,000 times3 and countless conference proceedings. These publications are publicly available through OSTI.gov and in the open literature. Additional scientific and technical products including project milestones that are not publications and NEUP project final reports are vetted through OSTI.gov and released once reviewed and approved by DOE. Since 2009, NEUP has awarded close to 1,000 different R&D projects in technical areas across the NE research programs. As of June 2023, 512 NEUP reports are publicly available on OSTI. The underlying data for projects is still held at universities, and data transfer, co-location, and dissemination has not occurred in a systematic way. NEUP data is currently accessible through myriad university-based data repositories, or through direct requests to PIs. The program identified this patchwork of repositories, or often lack of publicly available data, as a significant barrier to an organized, accessible, and comprehensive solution to sharing data with the larger nuclear energy community. Approach The goal of this pilot project is to establish a pathway to a consolidated long-term repository for NEUP project data. To accomplish this goal, the pilot strives to accomplish the following objectives: Establish data collection standards, including a standard set of required supplementary information to contextualize and support raw data files. Work with the HPC group collect and upload information and to modify the NRDS system, as needed, to support a standardized approach. Resolve potential barriers to successful roll out of an expanded data collection strategy, including modifying data management plan guidelines and establishing a document and data release process that accounts for potential intellectual property and/or export control concerns. Results Overall, the pilot was successful in collecting 8,982 raw and processes data files, 220 reports, 56 calibration files, and 5,931 other supplementary documents. Supplementary documents included experimental plans, methods, journal publications and conference proceedings, milestone reports, and final reports. Figure 2 shows the number of data sets and supplementary project information provided by each project. Projects has significantly different input, depending on experimental data produced and completeness of the datasets provided.

Data collection↗

The DECam Local Volume Exploration Survey: Overview and First Data Release

The DECam Local Volume Exploration survey (DELVE) is a 126-night survey program on the 4 m Blanco Telescope at the Cerro Tololo Inter-American Observatory in Chile. DELVE seeks to understand the characteristics of faint satellite galaxies and other resolved stellar substructures over a range of environments in the Local Volume. DELVE will combine new DECam observations with archival DECam data to cover ~15,000 deg 2 of high Galactic latitude (|b| > 10°) southern sky to a 5σ depth of g, r, i, z ~ 23.5 mag. In addition, DELVE will cover a region of ~2200 deg 2 around the Magellanic Clouds to a depth of g, r, i ~ 24.5 mag and an area of ~135 deg 2 around four Magellanic analogs to a depth of g, i ~ 25.5 mag. Here, we present an overview of the DELVE program and progress to date. Furthermore, we also summarize the first DELVE public data release (DELVE DR1), which provides point-source and automatic aperture photometry for ~520 million astronomical sources covering ~5000 deg 2 of the southern sky to a 5σ point-source depth of g = 24.3 mag, r = 23.9 mag, i = 23.3 mag, and z = 22.8 mag. DELVE DR1 is publicly available via the NOIRLab Astro Data Lab science platform.

79 ASTRONOMY AND ASTROPHYSICS↗

Journey to Time-Variable Moment Tensors through Inversion of Acoustic and Seismoacoustic Data

We explore the capability of acoustic and seismoacoustic datasets to directly resolve a complex, time-variable source consisting of a buried mechanism, represented as a moment tensor, and a spall mechanism, represented as a vertical force at the surface. Traditionally, each component of a resolved moment tensor assumes one underlying source time function, which likely fails to capture the full evolution of a dynamic source, such as an explosion followed by slip on near-source joints or development of spallation. Specifically, we expand previous work to resolve a time-variable moment tensor using single-modality and joint-modality inversion frameworks through analysis of infrasound and seismoacoustic data recorded as part of the Source Physics Experiment Phase II: Dry Alluvium Geology (DAG). We investigate the impact of including signals from seismic-to-air coupling that are local to each infrasound sensor in comparison to mainly atmosphere-propagating acoustic signals, which occur from coupling of the wavefield from the subsurface to the atmosphere directly above the source. Additionally, we assess the ability of our inversion algorithm to fit observed infrasound data using a variety of time-variable source mechanisms. First, we consider the buried moment tensor source alone, which assumes that the determined Green’s functions incorporate effects from spallation or that the impact from spallation is minimal. Second, we examine the estimated buried moment tensor and vertical surface spallation as terms that must both be resolved in the inversion. Third, we assess the ability for an estimated vertical surface spallation source to fit the acoustic data on its own. Finally, we compare results from the joint inversion of both seismic geophone and infrasound acoustic data for the buried-only source compared to buried and spallation sources. Our results are a preliminary investigation into the applications of the inversion technique to recorded datasets and show the technique has limited capabilities using acoustic data alone. Instead, this method shows promise for seismic and seismoacoustic datasets to resolve the time-variable mechanisms of a buried source.

47 OTHER INSTRUMENTATION↗