Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data exploration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Malicious Cyber Activity Detection using Zigzag Persistence

In this study we synthesize zigzag persistence from topological data analysis with autoencoder-based approaches to detect malicious cyber activity, and derive analytic insights. Cybersecurity aims to safeguard computers, networks, and servers from various forms of malicious attacks, including network damage, data theft, and activity monitoring. We focus on the cybersecurity domain and investigate the detection of malicious activity using log data. We consider the dynamics of the log data and explore the changing topology of a hypergraph representation of this data to gain insights into the underlying activity. These hypergraphs capture complex interactions between processes, together with their temporal information. To study the changing topology we use zigzag persistence, which captures how topological features persist at multiple dimensions over time. We observe that this detects malicious activity in a cyber data set. To automate this detection we implement an autoencoder trained on a vectorization of the resulting zigzag persistence barcodes. Our experimental results demonstrate the effectiveness of the autoencoder in detecting malicious activity. Overall, this study highlights the potential of zigzag persistence and its combination with temporal hypergraphs for analyzing cybersecurity log data and detecting malicious behavior.

hypergraphs, temporal hypergraph, topological data↗

Transforming Energy Through Computational Excellence: Advanced Scientific Visualization Reveals Energy Insights

The National Renewable Energy Laboratory's world-class researchers and analysts, along with the Insight Center (our state-of-the-art scientific visualization facility) make data immersion a reality, allowing users to step into and explore their data. With the rise of large, diverse, and distributed data sets, scientific visualization is now critical to the process of scientific discovery and to managing and analyzing data and extracting insights. NREL provides visualization capabilities and facilities that are supported by state-of-the-art equipment, leading-edge techniques, and expert staff.

data science↗

Uncertainty-Informed Volume Visualization using Implicit Neural Representation

The increasing adoption of Deep Neural Networks (DNNs) has led to their application in many challenging scientific visualization tasks. While advanced DNNs offer impressive generalization capabilities, understanding factors such as model prediction quality, robustness, and uncertainty is crucial. These insights can enable domain scientists to make informed decisions about their data. However, DNNs inherently lack ability to estimate prediction uncertainty, necessitating new research to construct robust uncertainty-aware visualization techniques tailored for various visualization tasks. In this work, we propose uncertainty-aware implicit neural representations to model scalar field data sets effectively and comprehensively study the efficacy and benefits of estimated uncertainty information for volume visualization tasks. We evaluate the effectiveness of two principled deep uncertainty estimation techniques: (1) Deep Ensemble and (2) Monte Carlo Dropout (MC-Dropout). These techniques enable uncertainty-informed volume visualization in scalar field data sets. Our extensive exploration across multiple data sets demonstrates that uncertainty-aware models produce informative volume visualization results. Moreover, integrating prediction uncertainty enhances the trustworthiness of our DNN model, making it suitable for robustly analyzing and visualizing real-world scientific volumetric data sets.

Saklani, Shanu↗

Modification and analysis of context-specific genome-scale metabolic models: methane-utilizing microbial chassis as a case study

ABSTRACT Context-specific genome-scale model (CS-GSM) reconstruction is becoming an efficient strategy for integrating and cross-comparing experimental multi-scale data to explore the relationship between cellular genotypes, facilitating fundamental or applied research discoveries. However, the application of CS modeling for non-conventional microbes is still challenging. Here, we present a graphical user interface that integrates COBRApy, EscherPy, and RIPTiDe, Python-based tools within the BioUML platform, and streamlines the reconstruction and interrogation of the CS genome-scale metabolic frameworks via Jupyter Notebook. The approach was tested using -omics data collected for Methylotuvimicrobium alcaliphilum 20Z R , a prominent microbial chassis for methane capturing and valorization. We optimized the previously reconstructed whole genome-scale metabolic network by adjusting the flux distribution using gene expression data. The outputs of the automatically reconstructed CS metabolic network were comparable to manually optimized i IA409 models for Ca-growth conditions. However, the CS model questions the reversibility of the phosphoketolase pathway and suggests higher flux via primary oxidation pathways. The model also highlighted unresolved carbon partitioning between assimilatory and catabolic pathways at the formaldehyde-formate node. Only a very few genes and only one enzyme with a predicted function in C1 metabolism, a homolog of the formaldehyde oxidation enzyme ( fae1-2 ), showed a significant change in expression in La-growth conditions. The CS-GSM predictions agreed with the experimental measurements under the assumption that the Fae1-2 is a part of the tetrahydrofolate-linked pathway. The cellular roles of the tungsten (W)-dependent formate dehydrogenase ( fdhAB ) and fae homologs ( fae1-2 and fae3 ) were investigated via mutagenesis. The phenotype of the f dhAB mutant followed the model prediction. Furthermore, a more significant reduction of the biomass yield was observed during growth in La-supplemented media, confirming a higher flux through formate. M. alcaliphilum 20Z R mutants lacking fae1-2 did not display any significant defects in methane or methanol-dependent growth. However, contrary to fae1, the fae1-2 homolog failed to restore the formaldehyde-activating enzyme function in complementation tests. Overall, the presented data suggest that the developed computational workflow supports the reconstruction and validation of CS-GSM networks of non-model microbes. IMPORTANCE The interrogation of various types of data is a routine strategy to explore the relationship between genotype and phenotype. An efficient approach for integrating and cross-comparing experimental multi-scale data in the context of whole-genome-based metabolic network reconstruction becomes a powerful tool that facilitates fundamental and applied research discoveries. The present study describes the reconstruction of a context-specific (CS) model for the methane-utilizing bacterium, Methylotuvimicrobium alcaliphilum 20Z R . M. alcaliphilum 20Z R is becoming an attractive microbial platform for the production of biofuels, chemicals, pharmaceuticals, and bio-sorbents for capturing atmospheric methane. We demonstrate that this pipeline can help reconstruct metabolic models that are similar to manually curated networks. Furthermore, the model is able to highlight previously overlooked pathways, thus advancing fundamental knowledge of non-model microbial systems or promoting their development toward biotechnological or environmental implementations.

Kulyashov, M. A.↗

FleetREDI Dashboard Fleet DNA Data Summaries

Developing daily duty cycle summaries for every vehicle-day within NLR’s Fleet DNA database was a key output of the FleetREDI project. This project captured second-by-second GPS and controller area network (CAN) data on in-use medium- and heavy-duty fleet vehicles and then summarized the data to provide an overview of vehicle operation throughout the United States. These data summaries were then displayed in aggregated formats on the FleetREDI dashboard, where users can explore the data within Fleet DNA. Fleet DNA’s clearinghouse of commercial fleet vehicle operating data helps vehicle manufacturers and developers optimize vehicle designs and helps fleet managers choose advanced technologies for their fleets. This online tool, which provides data summaries and visualizations similar to real-world "genetics" for medium- and heavy-duty fleet vehicles, helps users understand the broad operational range of commercial vehicles across vocations and weight classes.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Automated qualification data tool for high temperature metallic materials

This report describes a framework for storing, processing, and displaying qualification data for high temperature mechanical properties. The framework automates the process of generating design data from mechanical test results, for example for a data qualification report for the ASME Boiler \& Pressure Vessel Code. The framework has three parts: a data storage model with common formats for several types of typical mechanical property tests, a backend based on the \pycreep Python library for correlating and extrapolating the data to generate design material properties and allowable stresses, and a demonstration user interface for displaying, sorting, and filtering the data and exploring different options for modeling the design mechanical properties. The report discusses the options available for data processing, with illustrations from real test data on Alloy 617, Alloy 709, Alloy 740H, and Laser-Powder Bed Fusion 316H. The framework is complete for ASME type data analysis and will be used to store test data generated by the Department of Energy, Office of Nuclear Energy, Advanced Materials and Manufacturing Technologies sponsored qualification programs. Future work could extend the tool to other types of material properties and/or expand the demo user interface to make it accessible across the AMMT program.

36 MATERIALS SCIENCE↗

Approximate Dynamic Programming With Enhanced Off-Policy Learning for Coordinating Distributed Energy Resources

Herein this paper proposes an innovative approximate dynamic programming (ADP) method for distributed energy resource coordination with the loss of life of battery energy storage system (BESS) explicitly modeled. The dispatch policy is designed to account for both calendrical and cyclical aging effects on BESS, explicitly modeling the impacts of ambient temperature on BESS lifespan. The proposed ADP employs an adaptive critic method and enhanced off-policy deterministic policy gradient (DPG) strategy, addressing the limitations of the on-policy gradient-based ADP approaches, including inadequate exploration, low data usage, and computational complexity. In particular, a customized policy is proposed to guide the algorithm to explore some promising decisions and thereby improve exploration capability and learning efficiency compared to conventional DPG-based learning approaches, which may struggle to find a global optimum due to random noisy action-based exploration or require expert demonstration with extra effort. The proposed method is illustrated using the IEEE 123-node system and compared with the existing ADP methods to prove solution accuracy and demonstrate the effects of incorporating degradation models into control design. Case studies showed that the proposed ADP effectively coordinates DERs with a 10 times smaller optimization gap compared to existing methods, and the incorporation of the BESS life loss model ensures the expected lifespan and results in significant cost savings.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗

Expanding Access to Science Participation: A FAIR Framework for Petascale Data Visualization and Analytics

The massive data generated by scientists daily serve as both a major catalyst for new discoveries and innovations, as well as a significant roadblock that restricts access to the data. Here, our paper introduces a new approach to removing Big Data barriers and democratizing access to petascale data for the broader scientific community. Our novel data fabric abstraction layer allows user-friendly querying of scientific information while hiding the complexities of dealing with file systems or cloud services. We enable FAIR (Findable, Accessible, Interoperable, and Reusable) access to datasets such as NASA’s petascale climate datasets. Our paper presents an approach to managing, visualizing, and analyzing petabytes of data within a browser on equipment ranging from the top NASA supercomputer to commodity hardware like a laptop. Our novel data fabric abstraction utilizes state-of-the art progressive compression algorithms and machine-learning insights to power scalable visualization dashboards for petascale data. The result provides users with the ability to identify extreme events or trends dynamically, expanding access to scientific data and further enabling discoveries. We validate our approach by improving the ability of climate scientists to visually explore their data via three fully interactive dashboards. We further validate our approach by deploying the dashboards and simplified training materials in the classroom at a minority-serving institution. These dashboards, released in simplified form to the general public, contribute significantly to a broader push to democratize the access and use of climate data.

Computer science↗

Novel, active, and uncultured hydrocarbon-degrading microbes in the ocean

ABSTRACT Given the vast quantity of oil and gas input to the marine environment annually, hydrocarbon degradation by marine microorganisms is an essential ecosystem service. Linkages between taxonomy and hydrocarbon degradation capabilities are largely based on cultivation studies, leaving a knowledge gap regarding the intrinsic ability of uncultured marine microbes to degrade hydrocarbons. To address this knowledge gap, metagenomic sequence data from the Deepwater Horizon (DWH) oil spill deep-sea plume was assembled to which metagenomic and metatranscriptomic reads were mapped. Assembly and binning produced new DWH metagenome-assembled genomes that were evaluated along with their close relatives, all of which are from the marine environment (38 total). These analyses revealed globally distributed hydrocarbon-degrading microbes with clade-specific substrate degradation potentials that have not been reported previously. For example, methane oxidation capabilities were identified in all Cycloclasticus . Furthermore, all Bermanella encoded and expressed genes for non-gaseous n -alkane degradation; however, DWH Bermanella encoded alkane hydroxylase, not alkane 1-monooxygenase. All but one previously unrecognized DWH plume member in the SAR324 and UBA11654 have the capacity for aromatic hydrocarbon degradation. In contrast, Colwellia were diverse in the hydrocarbon substrates they could degrade. All clades encoded nutrient acquisition strategies and response to cold temperatures, while sensory and acquisition capabilities were clade specific. These novel insights regarding hydrocarbon degradation by uncultured planktonic microbes provides missing data, allowing for better prediction of the fate of oil and gas when hydrocarbons are input to the ocean, leading to a greater understanding of the ecological consequences to the marine environment. IMPORTANCE Microbial degradation of hydrocarbons is a critically important process promoting ecosystem health, yet much of what is known about this process is based on physiological experiments with a few hydrocarbon substrates and cultured microbes. Thus, the ability to degrade the diversity of hydrocarbons that comprise oil and gas by microbes in the environment, particularly in the ocean, is not well characterized. Therefore, this study aimed to utilize non-cultivation-based ‘omics data to explore novel genomes of uncultured marine microbes involved in degradation of oil and gas. Analyses of newly assembled metagenomic data and previously existing genomes from other marine data sets, with metagenomic and metatranscriptomic read recruitment, revealed globally distributed hydrocarbon-degrading marine microbes with clade-specific substrate degradation potentials that have not been previously reported. This new understanding of oil and gas degradation by uncultured marine microbes suggested that the global ocean harbors a diversity of hydrocarbon-degrading bacteria, which can act as primary agents regulating ecosystem health.

Howe, Kathryn L.↗

“Translational Opportunities in CPS Transportation”

This talk will describe opportunities for translating research from open-road experiments with modified adaptive cruise controllers. The data and controllers used for the previous research are based on use cases and example drives from within the US. We explore processes, data sharing, experiments, and other techniques that we think will drive trans-Pacific partnerships with driving data.

Sprinkle, Jonathan↗

Model Choice Metrics to Optimize Profile-QSAR Performance

Predicting molecular activity against protein targets is difficult because of the paucity of experimental data. Approaches like multitask modeling and collaborative filtering seek to improve model accuracy by leveraging results from multiple targets, but are limited because different compounds are measured with different assays, leading to sparse data matrices. Profile-QSAR (pQSAR) 2.0 addresses this problem by fitting a series of partial least squares models for each target, using as features the predictions from single-task models on the remaining targets. Here, this method has been shown to produce better results than single task and multitask models. However, the factors determining the success of pQSAR 2.0 have as yet not been characterized. In this paper we examine the experimental conditions that lead to better pQSAR models. We limit the amount of data available to the method by retraining with decreasing amounts of data and explore the model’s ability to generalize to compounds that have never been assayed. Finally, we look at the properties of training data needed to demonstrate pQSAR improvement.

Biological and medical sciences, Computer science↗

Integrating Data Centers and Grid Technologies at Scale

This presentation focuses on the challenge of integrating AI-driven data centers with the power grid at scale. It examines the AI data center capacity challenge and the role of new Medium Voltage Direct Current (MVDC) and other grid-enhancing technologies in enabling efficient and reliable power delivery. The session will highlight the National Laboratory of the Rockies' ARIES capabilities and planning tools, along with collaborative examples involving Verrus, Compass, and Schneider through the Agora test bed for grid-friendly data center evaluations, and ON. Energy for UPS evaluation. It will showcase the NLR Stable Grid Platform for studying oscillations caused by large-scale data centers, along with planning tools to assess grid security and reliability. Additionally, the presentation covers reconductoring strategies to increase grid capacity and explores innovative data center architectures, including the Advanced DC Architectures with Power-electronic Transformers (ADAPT) platform, which enables testing of complete DC architectures for data centers.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Improved Earthquake Source Parameters with 3D Wavespeed Models in California and Nevada

Seismic tomography harnesses earthquake data to explore the inaccessible structure of the Earth. Adjoint waveform tomography (AWT), a method of seismic tomography, updates the tomographic model by optimizing the fit between observed earthquake data and synthetic waveforms. The synthetic data are calculated by solving the wave equation through a given 3D model. An important requirement to calculating synthetics is the source information (location, centroid time, depth, and moment tensor). Errors in source information affect the quality of the synthetics produced, which in turn can limit how structure can be inferred in the AWT workflow. Here, to test the effect of updating source information, we used MTTime (Chiang, 2020), a time-domain full-waveform moment tensor inversion code, to calculate the moment tensors and depths of 118 earthquakes that occurred in California and Nevada over a 20-yr period. We calculated 3D Green’s functions using a 3D seismic wavespeed model of California and Nevada (Doody et al., 2023b). We show that the inverted solutions provide better waveform fits than the Global Centroid Moment Tensor catalog and increase usable, well-correlated data by up to 7%. Therefore, we argue that recalculating source parameters should be considered in AWT workflows, particularly for smaller magnitude events (⁠M w > 5.0).

58 GEOSCIENCES↗

Hawai‘i Supernova Flows: a peculiar velocity survey using over a Thousand Supernovae in the near-infrared

ABSTRACT We introduce the Hawai‘i Supernova Flows project and present summary statistics of the first 1217 astronomical transients observed, 668 of which are spectroscopically classified Type Ia Supernovae (SNe Ia). Our project is designed to obtain systematics-limited distances to SNe Ia while consuming minimal dedicated observational resources. To date, we have performed almost 5000 near-infrared (NIR) observations of astronomical transients and have obtained spectra for over 200 host galaxies lacking published spectroscopic redshifts. In this survey paper, we describe the methodology used to select targets, collect/reduce data, calculate distances, and perform quality cuts. We compare our methods to those used in similar studies, finding general agreement or mild improvement. Our summary statistics include various parametrizations of dispersion in the Hubble diagrams produced using fits to several commonly used SN Ia models. We find the lowest dispersions using the SNooPy package’s EBV_model2, with a root mean square deviation of 0.165 mag and a normalized median absolute deviation of 0.123 mag. The full utility of the Hawai‘i Supernova Flows data set far exceeds the analyses presented in this paper. Our photometry will provide a valuable test bed for models of SN Ia incorporating NIR data. Differential cosmological studies comparing optical samples and combined optical and NIR samples will have increased leverage for constraining chromatic effects like dust extinction. We invite the community to explore our data by making the light curves, fits, and host galaxy redshifts publicly accessible.

Do, Aaron (ORCID:0000000334297845)↗

Phase Selection Rules of Multi‐Principal Element Alloys

Abstract Computational prediction of phase stability of multi‐principal element alloys (MPEAs) holds a lot of promise for rapid exploration of the enormous design space and autonomous discovery of superior structural and functional properties. Regardless of many plausible works that rely on phenomenological theory and machine learning, precise prediction is still limited by insufficient data and the lack of interpretability of some machine learning algorithms, e.g., convolutional neural network. In this work, a comprehensive approach is presented, encompassing the development of a complete dataset that contains 72 387 density functional theory calculations, as well as a predictive global phenomenological descriptor. The phase selection descriptor, based on atomic electronegativity and valence electron concentration, significantly outperforms the widely used valence electron concentration, excelling in both accuracy (with an f1 score of 63% compared to 47%) and its ability to predict the HCP phase (0.48 recall compared to 0). The comprehensive data mining on the global design space of 61 425 quaternary MPEAs made from 28 possible metals, together with the phenomenological theory and physical interpretation, will set up a solid computational science foundation for data‐driven exploration of MPEAs.

Chemistry↗

Cosmological preference for a negative neutrino mass

The most precise determination of the sum of neutrino masses from cosmological data, derived from analysis of the cosmic microwave background (CMB) and baryon acoustic acoustic oscillations (BAO) from the Dark Energy Spectroscopic Instrument (DESI), favors a value below the minimum inferred from neutrino flavor oscillation experiments. We explore which data is most responsible of this puzzling aspect of the current constraints on neutrino mass and whether it is related to other anomalies in cosmology. We demonstrate conclusively that the preference for negative neutrino masses is a consequence of larger than expected lensing of the CMB in both the two- and four-point lensing statistics. Furthermore, we show that this preference is robust to changes in likelihoods of the BAO and CMB optical depth analyses given the available data. We then show that this excess clustering is not easily explained by changes to the expansion history and is likely distinct from the preference for for dynamical dark energy in DESI BAO data. Finally, we discuss how future data may impact these results, including an analysis of Planck CMB with mock DESI 5-year data. Here, we conclude that the negative neutrino mass preference is likely to persist even as more cosmological data is collected in the near future.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗