Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data exploration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Performance Debugging and Tuning of Flash-X with Data Analysis Tools

State-of-the-art multiphysics simulations running on large scale leadership computing platforms have many variables contributing to their performance and scaling behavior. We recently encountered an interesting performance anomaly in Flash-X, a multiphysics multicomponent simulation software, when characterizing its performance behavior on several large-scale HPC platforms. The anomaly was tracked down to the interaction between the use of dynamic allocation of scratch data and data locality in the cache hierarchy. In this paper we present the details of unexpected performance variability of Flash-X, its extensive analysis using the performance measurement tool TAU to collect the data and Python data analysis libraries to explore the data, and our insights from this experience. In this process, we discovered and removed or mitigated two additional performance limiting bottlenecks for performance tuning.

Huck, Kevin↗

University Data Management Pilot Utilizing the Nuclear Research Data System

Background In 2022, the Office of Science and Technology Policy (OSTP) issued a memo that significantly reshaped the landscape of access to federally funded research. The memo mandated that all taxpayer-funded research be made available to the public without delay upon publication, without an embargo period, superseding the 2013 OSTP public access policy. This public access policy promotes transparency and the democratization of knowledge, ensuring that the fruits of scientific endeavors funded by federal agencies could be immediately accessed and built upon by scientists, educators, students, and the public at large. To implement the requirements of the OSTP guidance and DOE Public Access Plan, the Office of Nuclear Energy (NE) has implemented public access plan guidance and has identified several areas where better data management practices would further expand public access to important nuclear energy related scientific data, reports, and other technical products. Significant NE supported efforts are already underway for data management and public access to important nuclear energy related data.1 2 To address gaps in data management practices, and improve retention and accessibility of data, NE is actively exploring enhanced data management options utilizing its high-performance computing resources administered by its Nuclear Scientific User Facility Program. A newly piloted system, the Nuclear Research Data System (NRDS) acts as a portal for data collection and dissemination. Nuclear Energy University Program Research and Development Portfolio According to Web of Science, NEUP has produced 2,345 journal publication that have been cited more than 61,000 times3 and countless conference proceedings. These publications are publicly available through OSTI.gov and in the open literature. Additional scientific and technical products including project milestones that are not publications and NEUP project final reports are vetted through OSTI.gov and released once reviewed and approved by DOE. Since 2009, NEUP has awarded close to 1,000 different R&D projects in technical areas across the NE research programs. As of June 2023, 512 NEUP reports are publicly available on OSTI. The underlying data for projects is still held at universities, and data transfer, co-location, and dissemination has not occurred in a systematic way. NEUP data is currently accessible through myriad university-based data repositories, or through direct requests to PIs. The program identified this patchwork of repositories, or often lack of publicly available data, as a significant barrier to an organized, accessible, and comprehensive solution to sharing data with the larger nuclear energy community. Approach The goal of this pilot project is to establish a pathway to a consolidated long-term repository for NEUP project data. To accomplish this goal, the pilot strives to accomplish the following objectives: Establish data collection standards, including a standard set of required supplementary information to contextualize and support raw data files. Work with the HPC group collect and upload information and to modify the NRDS system, as needed, to support a standardized approach. Resolve potential barriers to successful roll out of an expanded data collection strategy, including modifying data management plan guidelines and establishing a document and data release process that accounts for potential intellectual property and/or export control concerns. Results Overall, the pilot was successful in collecting 8,982 raw and processes data files, 220 reports, 56 calibration files, and 5,931 other supplementary documents. Supplementary documents included experimental plans, methods, journal publications and conference proceedings, milestone reports, and final reports. Figure 2 shows the number of data sets and supplementary project information provided by each project. Projects has significantly different input, depending on experimental data produced and completeness of the datasets provided.

Data collection↗

The DECam Local Volume Exploration Survey: Overview and First Data Release

The DECam Local Volume Exploration survey (DELVE) is a 126-night survey program on the 4 m Blanco Telescope at the Cerro Tololo Inter-American Observatory in Chile. DELVE seeks to understand the characteristics of faint satellite galaxies and other resolved stellar substructures over a range of environments in the Local Volume. DELVE will combine new DECam observations with archival DECam data to cover ~15,000 deg 2 of high Galactic latitude (|b| > 10°) southern sky to a 5σ depth of g, r, i, z ~ 23.5 mag. In addition, DELVE will cover a region of ~2200 deg 2 around the Magellanic Clouds to a depth of g, r, i ~ 24.5 mag and an area of ~135 deg 2 around four Magellanic analogs to a depth of g, i ~ 25.5 mag. Here, we present an overview of the DELVE program and progress to date. Furthermore, we also summarize the first DELVE public data release (DELVE DR1), which provides point-source and automatic aperture photometry for ~520 million astronomical sources covering ~5000 deg 2 of the southern sky to a 5σ point-source depth of g = 24.3 mag, r = 23.9 mag, i = 23.3 mag, and z = 22.8 mag. DELVE DR1 is publicly available via the NOIRLab Astro Data Lab science platform.

79 ASTRONOMY AND ASTROPHYSICS↗

Journey to Time-Variable Moment Tensors through Inversion of Acoustic and Seismoacoustic Data

We explore the capability of acoustic and seismoacoustic datasets to directly resolve a complex, time-variable source consisting of a buried mechanism, represented as a moment tensor, and a spall mechanism, represented as a vertical force at the surface. Traditionally, each component of a resolved moment tensor assumes one underlying source time function, which likely fails to capture the full evolution of a dynamic source, such as an explosion followed by slip on near-source joints or development of spallation. Specifically, we expand previous work to resolve a time-variable moment tensor using single-modality and joint-modality inversion frameworks through analysis of infrasound and seismoacoustic data recorded as part of the Source Physics Experiment Phase II: Dry Alluvium Geology (DAG). We investigate the impact of including signals from seismic-to-air coupling that are local to each infrasound sensor in comparison to mainly atmosphere-propagating acoustic signals, which occur from coupling of the wavefield from the subsurface to the atmosphere directly above the source. Additionally, we assess the ability of our inversion algorithm to fit observed infrasound data using a variety of time-variable source mechanisms. First, we consider the buried moment tensor source alone, which assumes that the determined Green’s functions incorporate effects from spallation or that the impact from spallation is minimal. Second, we examine the estimated buried moment tensor and vertical surface spallation as terms that must both be resolved in the inversion. Third, we assess the ability for an estimated vertical surface spallation source to fit the acoustic data on its own. Finally, we compare results from the joint inversion of both seismic geophone and infrasound acoustic data for the buried-only source compared to buried and spallation sources. Our results are a preliminary investigation into the applications of the inversion technique to recorded datasets and show the technique has limited capabilities using acoustic data alone. Instead, this method shows promise for seismic and seismoacoustic datasets to resolve the time-variable mechanisms of a buried source.

47 OTHER INSTRUMENTATION↗

Assessing complexity and dynamics in epidemics: geographical barriers and facilitators of foot-and-mouth disease dissemination

Introduction: Physical and non-physical processes that occur in nature may influence biological processes, such as dissemination of infectious diseases. However, such processes may be hard to detect when they are complex systems. Because complexity is a dynamic and non-linear interaction among numerous elements and structural levels in which specific effects are not necessarily linked to any one specific element, cause-effect connections are rarely or poorly observed. Methods: To test this hypothesis, the complex and dynamic properties of geo-biological data were explored with high-resolution epidemiological data collected in the 2001 Uruguayan foot-and-mouth disease (FMD) epizootic that mainly affected cattle. County-level data on cases, farm density, road density, river density, and the ratio of road (or river) length/county perimeter were analyzed with an open-ended procedure that identified geographical clustering in the first 11 epidemic weeks. Two questions were asked: (i) do geo-referenced epidemiologic data display complex properties? and (ii) can such properties facilitate or prevent disease dissemination? Results: Emergent patterns were detected when complex data structures were analyzed, which were not observed when variables were assessed individually. Complex properties–including data circularity–were demonstrated. The emergent patterns helped identify 11 counties as ‘disseminators’ or ‘facilitators’ (F) and 264 counties as ‘barriers’ (B) of epidemic spread. In the early epidemic phase, F and B counties differed in terms of road density and FMD case density. Focusing on non-biological, geographical data, a second analysis indicated that complex relationships may identify B-like counties even before epidemics occur. Discussion: Geographical barriers and/or promoters of disease dispersal may precede the introduction of emerging pathogens. If corroborated, the analysis of geo-referenced complexity may support anticipatory epidemiological policies.

60 APPLIED LIFE SCIENCES↗

Data and Tools for Exploring New Pumped Storage Hydropower Deployment Opportunities

Pumped storage hydropower (PSH) is a flexible energy storage technology with the potential to facilitate variable renewable energy integration into the decarbonized electric grid of the future. NREL is developing new data and tools to help understand opportunities for new PSH deployment, including nationwide resource assessment data, a bottom-up component-level cost model, and a lifecycle greenhouse gas emissions calculator. These datasets lay the foundation for better-informed grid planning decisions about how PSH fits into a future portfolio of generation, transmission, and storage assets.

CEM↗

A Practical Comparison of Data-Driven Prognostics Methods for Energy Systems

This study explores data-driven prognostics for nuclear power plant (NPP) condensers, focusing on tube fouling. We utilized the Asherah nuclear power plant simulator (ANS) to compare four methods: Random Forest (RF), Support Vector Regressor (SVR), Fully Connected Neural Network (FCNN), and Long Short-Term Memory Neural Network (LSTM). By simulating various fouling scenarios in the ANS, we generated data with different degradation rates under transient operations. The models were trained and tested on these data, with performance evaluated visually and numerically including uncertainty assessment. The LSTM model excelled, exhibiting minimal prediction noise and the most accurate remaining useful life estimates across all degradation levels. Its ability to capture long-term dependencies and produce cleaner outputs makes it a strong candidate, although accurate training data across the entire component lifespan are crucial. The RF model emerged as a robust alternative, providing reliable predictions with high confidence. The FCNN and SVR models, while less effective overall, showed potential under specific conditions. FCNN offers a less complex alternative to LSTM and might benefit from larger datasets. SVR excels in precision when the quality of the training data is high. Furthermore, this study highlights the operational benefits of advanced prognostics in the energy sector and emphasizes the need for further research in NPP condenser health management through real-life experiments.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

A VOI Web Application for Distinct Geothermal Domains: Statistical Evaluation of Different Data Types within the Great Basin

The Great Basin region contains different domains that have different structural and hydrothermal flow patterns. Depending on the characteristics of these patterns, certain data types may be more successful at detecting hidden geothermal resources. In this paper, we quantitatively evaluate if certain data types are more successful in certain domains. Given different aquifer, strain and structural conditions, we explore which data types statistically reveal positively labeled geothermal sites. We utilize value of information (VOI) metrics to help quantify the reliability of data types to discriminate against "positive" and "negative" labeled geothermal sites. We also evaluate how kernel density estimation can help generalize the statistics that inform VOI, which is necessary given the limited data in geothermal exploration. Except for the Carbonate Aquifer, the highest ranking of the Vimperfect is the Local Structural Setting. Next, the slip and dilation tendency is first for Carbonate Aquifer and second for Central Nevada Seismic Belt and Western Great Basin. For the Carbonate Aquifer, heat flow is has the lowest Vimperfect value compared to the other three domains, which is consistent with the understanding of how heat flow measurements are masked by regional groundwater flow.

Bayesian analysis↗

Radioisotope Identification with List-Mode Gamma-Ray Data

This work explores the potential of utilizing temporal data from gamma-ray detectors, known as list-mode data, to enhance radioisotope identification. Traditional identification methods, which rely on full gamma-ray spectrum analysis, often require long dwell times and struggle with spectra containing similarly spaced spectral peaks. We hypothesize that by leveraging the probabilistic nature of nuclear decay and the time-encoded information from decay sequences and interactions with surrounding materials, we can improve classification accuracy over static spectral analysis. This research examines the temporal content of list-mode data through exploratory data analysis via correlation discovery and qualitative distribution analysis. Additionally, we propose a probabilistic classification model that can utilize spectral data, temporal data, or both to determine if the incorporation of temporal information improves radioisotope identification. Our findings suggest that the temporal information present in list-mode gamma-ray data has merit and should be further investigated to develop more robust and optimal methods for utilizing this temporal information in applications requiring radioisotope identification.

List-mode data↗

Design-Space Data: Informing Common Design Decisions with Pre-Simulated Data

Design Space Exploration (DSE) analysis techniques represent a data-centric approach to integrating performance analysis in early design phases when there is the greatest potential to cheaply improve the energy efficiency of a building. We focus on a novel extension of DSE called Universal Design Space Exploration (UDSE), which leverages massive databases of pre-simulated analysis that represent all possible outcomes of common analysis workflows. These databases, called Design Spaces, become “universal” when a single pre-simulated design space can be re-applied to future unknown projects. Unlike current simulation methods, which require a design to exist before it can be analyzed and often take minutes or hours to simulate, UDSE leverages pre-simulation to deliver rapid and relevant insight as new designs are conceptualized. The data underpinning UDSE enables advanced statistical and Artificial Intelligence methods, allowing UDSE to deliver a greater understanding of the larger problem being explored, rather than simply delivering analysis of several pre-conceived design options. We believe that UDSE can provide instantaneous, relevant analysis for all building design projects at negligible cost. This paper has two main goals, to develop a relevant Universal Design Space that showcases the potential of UDSE and to release this data freely to industry and academia; thereby lowering the barrier to entry to digital literacy in statistics, ML and AI within the architecture, engineering, construction (AEC) industry.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

GNPS Dashboard: collaborative exploration of mass spectrometry data in the web browser

Access to web-based platforms has enabled scientists to perform research remotely. A critical aspect of mass spectrometry data analysis is the inspection, analysis, and visualization of the raw data to validate data quality and confirm statistical observations. We developed the GNPS Dashboard, a web-based data visualization tool, to facilitate synchronous collaborative inspection, visualization, and analysis of private and public mass spectrometry data remotely.

59 BASIC BIOLOGICAL SCIENCES↗

DeepUQ: Assessing the Aleatoric Uncertainties from two Deep Learning Methods

Assessing the quality of aleatoric uncertainty estimates from uncertainty quantification (UQ) deep learning methods is important in scientific contexts, where uncertainty is physically meaningful and important to characterize and interpret exactly. We systematically compare aleatoric uncertainty measured by two UQ techniques, Deep Ensembles (DE) and Deep Evidential Regression (DER). Our method focuses on both zero-dimensional (0D) and two-dimensional (2D) data, to explore how the UQ methods function for different data dimensionalities. We investigate uncertainty injected on the input and output variables and include a method to propagate uncertainty in the case of input uncertainty so that we can compare the predicted aleatoric uncertainty to the known values. We experiment with three levels of noise. The aleatoric uncertainty predicted across all models and experiments scales with the injected noise level. However, the predicted uncertainty is miscalibrated to $\rm{std}(\sigma_{\rm al})$ with the true uncertainty for half of the DE experiments and almost all of the DER experiments. The predicted uncertainty is the least accurate for both UQ methods for the 2D input uncertainty experiment and the high-noise level. While these results do not apply to more complex data, they highlight that further research on post-facto calibration for these methods would be beneficial, particularly for high-noise and high-dimensional settings.

Nevin, Rebecca↗

Internship Final Report on the unsupervised learning sensor fusion (ULSF) approach

This paper describes a summer internship project undertaken at Sandia National Labs (SNL), both current status and future work. The project was to explore various machine learning approaches for use on turbulent flow data. Specifically, unsupervised classification of turbulent flow data was explored. First, the usage of models in this field is discussed, and several issues in the common usage of the models are identified. Solutions to these issues are then proposed, in the form of a Bayesian filtering approach which probabilistically incorporates multiple sources of data to improve confidence in a result. Several types of sensors are suggested for this method, the incorporation of which range from semi-supervised learning approaches to fully unsupervised. These approaches are then tested on several turbulent flow cases.

97 MATHEMATICS AND COMPUTING↗

Malicious Cyber Activity Detection using Zigzag Persistence

In this study we synthesize zigzag persistence from topological data analysis with autoencoder-based approaches to detect malicious cyber activity, and derive analytic insights. Cybersecurity aims to safeguard computers, networks, and servers from various forms of malicious attacks, including network damage, data theft, and activity monitoring. We focus on the cybersecurity domain and investigate the detection of malicious activity using log data. We consider the dynamics of the log data and explore the changing topology of a hypergraph representation of this data to gain insights into the underlying activity. These hypergraphs capture complex interactions between processes, together with their temporal information. To study the changing topology we use zigzag persistence, which captures how topological features persist at multiple dimensions over time. We observe that this detects malicious activity in a cyber data set. To automate this detection we implement an autoencoder trained on a vectorization of the resulting zigzag persistence barcodes. Our experimental results demonstrate the effectiveness of the autoencoder in detecting malicious activity. Overall, this study highlights the potential of zigzag persistence and its combination with temporal hypergraphs for analyzing cybersecurity log data and detecting malicious behavior.

hypergraphs, temporal hypergraph, topological data↗

Transforming Energy Through Computational Excellence: Advanced Scientific Visualization Reveals Energy Insights

The National Renewable Energy Laboratory's world-class researchers and analysts, along with the Insight Center (our state-of-the-art scientific visualization facility) make data immersion a reality, allowing users to step into and explore their data. With the rise of large, diverse, and distributed data sets, scientific visualization is now critical to the process of scientific discovery and to managing and analyzing data and extracting insights. NREL provides visualization capabilities and facilities that are supported by state-of-the-art equipment, leading-edge techniques, and expert staff.

data science↗

Uncertainty-Informed Volume Visualization using Implicit Neural Representation

The increasing adoption of Deep Neural Networks (DNNs) has led to their application in many challenging scientific visualization tasks. While advanced DNNs offer impressive generalization capabilities, understanding factors such as model prediction quality, robustness, and uncertainty is crucial. These insights can enable domain scientists to make informed decisions about their data. However, DNNs inherently lack ability to estimate prediction uncertainty, necessitating new research to construct robust uncertainty-aware visualization techniques tailored for various visualization tasks. In this work, we propose uncertainty-aware implicit neural representations to model scalar field data sets effectively and comprehensively study the efficacy and benefits of estimated uncertainty information for volume visualization tasks. We evaluate the effectiveness of two principled deep uncertainty estimation techniques: (1) Deep Ensemble and (2) Monte Carlo Dropout (MC-Dropout). These techniques enable uncertainty-informed volume visualization in scalar field data sets. Our extensive exploration across multiple data sets demonstrates that uncertainty-aware models produce informative volume visualization results. Moreover, integrating prediction uncertainty enhances the trustworthiness of our DNN model, making it suitable for robustly analyzing and visualizing real-world scientific volumetric data sets.

Saklani, Shanu↗

Modification and analysis of context-specific genome-scale metabolic models: methane-utilizing microbial chassis as a case study

ABSTRACT Context-specific genome-scale model (CS-GSM) reconstruction is becoming an efficient strategy for integrating and cross-comparing experimental multi-scale data to explore the relationship between cellular genotypes, facilitating fundamental or applied research discoveries. However, the application of CS modeling for non-conventional microbes is still challenging. Here, we present a graphical user interface that integrates COBRApy, EscherPy, and RIPTiDe, Python-based tools within the BioUML platform, and streamlines the reconstruction and interrogation of the CS genome-scale metabolic frameworks via Jupyter Notebook. The approach was tested using -omics data collected for Methylotuvimicrobium alcaliphilum 20Z R , a prominent microbial chassis for methane capturing and valorization. We optimized the previously reconstructed whole genome-scale metabolic network by adjusting the flux distribution using gene expression data. The outputs of the automatically reconstructed CS metabolic network were comparable to manually optimized i IA409 models for Ca-growth conditions. However, the CS model questions the reversibility of the phosphoketolase pathway and suggests higher flux via primary oxidation pathways. The model also highlighted unresolved carbon partitioning between assimilatory and catabolic pathways at the formaldehyde-formate node. Only a very few genes and only one enzyme with a predicted function in C1 metabolism, a homolog of the formaldehyde oxidation enzyme ( fae1-2 ), showed a significant change in expression in La-growth conditions. The CS-GSM predictions agreed with the experimental measurements under the assumption that the Fae1-2 is a part of the tetrahydrofolate-linked pathway. The cellular roles of the tungsten (W)-dependent formate dehydrogenase ( fdhAB ) and fae homologs ( fae1-2 and fae3 ) were investigated via mutagenesis. The phenotype of the f dhAB mutant followed the model prediction. Furthermore, a more significant reduction of the biomass yield was observed during growth in La-supplemented media, confirming a higher flux through formate. M. alcaliphilum 20Z R mutants lacking fae1-2 did not display any significant defects in methane or methanol-dependent growth. However, contrary to fae1, the fae1-2 homolog failed to restore the formaldehyde-activating enzyme function in complementation tests. Overall, the presented data suggest that the developed computational workflow supports the reconstruction and validation of CS-GSM networks of non-model microbes. IMPORTANCE The interrogation of various types of data is a routine strategy to explore the relationship between genotype and phenotype. An efficient approach for integrating and cross-comparing experimental multi-scale data in the context of whole-genome-based metabolic network reconstruction becomes a powerful tool that facilitates fundamental and applied research discoveries. The present study describes the reconstruction of a context-specific (CS) model for the methane-utilizing bacterium, Methylotuvimicrobium alcaliphilum 20Z R . M. alcaliphilum 20Z R is becoming an attractive microbial platform for the production of biofuels, chemicals, pharmaceuticals, and bio-sorbents for capturing atmospheric methane. We demonstrate that this pipeline can help reconstruct metabolic models that are similar to manually curated networks. Furthermore, the model is able to highlight previously overlooked pathways, thus advancing fundamental knowledge of non-model microbial systems or promoting their development toward biotechnological or environmental implementations.

Kulyashov, M. A.↗