Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data exploration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗

Expanding Access to Science Participation: A FAIR Framework for Petascale Data Visualization and Analytics

The massive data generated by scientists daily serve as both a major catalyst for new discoveries and innovations, as well as a significant roadblock that restricts access to the data. Here, our paper introduces a new approach to removing Big Data barriers and democratizing access to petascale data for the broader scientific community. Our novel data fabric abstraction layer allows user-friendly querying of scientific information while hiding the complexities of dealing with file systems or cloud services. We enable FAIR (Findable, Accessible, Interoperable, and Reusable) access to datasets such as NASA’s petascale climate datasets. Our paper presents an approach to managing, visualizing, and analyzing petabytes of data within a browser on equipment ranging from the top NASA supercomputer to commodity hardware like a laptop. Our novel data fabric abstraction utilizes state-of-the art progressive compression algorithms and machine-learning insights to power scalable visualization dashboards for petascale data. The result provides users with the ability to identify extreme events or trends dynamically, expanding access to scientific data and further enabling discoveries. We validate our approach by improving the ability of climate scientists to visually explore their data via three fully interactive dashboards. We further validate our approach by deploying the dashboards and simplified training materials in the classroom at a minority-serving institution. These dashboards, released in simplified form to the general public, contribute significantly to a broader push to democratize the access and use of climate data.

Computer science↗

Novel, active, and uncultured hydrocarbon-degrading microbes in the ocean

ABSTRACT Given the vast quantity of oil and gas input to the marine environment annually, hydrocarbon degradation by marine microorganisms is an essential ecosystem service. Linkages between taxonomy and hydrocarbon degradation capabilities are largely based on cultivation studies, leaving a knowledge gap regarding the intrinsic ability of uncultured marine microbes to degrade hydrocarbons. To address this knowledge gap, metagenomic sequence data from the Deepwater Horizon (DWH) oil spill deep-sea plume was assembled to which metagenomic and metatranscriptomic reads were mapped. Assembly and binning produced new DWH metagenome-assembled genomes that were evaluated along with their close relatives, all of which are from the marine environment (38 total). These analyses revealed globally distributed hydrocarbon-degrading microbes with clade-specific substrate degradation potentials that have not been reported previously. For example, methane oxidation capabilities were identified in all Cycloclasticus . Furthermore, all Bermanella encoded and expressed genes for non-gaseous n -alkane degradation; however, DWH Bermanella encoded alkane hydroxylase, not alkane 1-monooxygenase. All but one previously unrecognized DWH plume member in the SAR324 and UBA11654 have the capacity for aromatic hydrocarbon degradation. In contrast, Colwellia were diverse in the hydrocarbon substrates they could degrade. All clades encoded nutrient acquisition strategies and response to cold temperatures, while sensory and acquisition capabilities were clade specific. These novel insights regarding hydrocarbon degradation by uncultured planktonic microbes provides missing data, allowing for better prediction of the fate of oil and gas when hydrocarbons are input to the ocean, leading to a greater understanding of the ecological consequences to the marine environment. IMPORTANCE Microbial degradation of hydrocarbons is a critically important process promoting ecosystem health, yet much of what is known about this process is based on physiological experiments with a few hydrocarbon substrates and cultured microbes. Thus, the ability to degrade the diversity of hydrocarbons that comprise oil and gas by microbes in the environment, particularly in the ocean, is not well characterized. Therefore, this study aimed to utilize non-cultivation-based ‘omics data to explore novel genomes of uncultured marine microbes involved in degradation of oil and gas. Analyses of newly assembled metagenomic data and previously existing genomes from other marine data sets, with metagenomic and metatranscriptomic read recruitment, revealed globally distributed hydrocarbon-degrading marine microbes with clade-specific substrate degradation potentials that have not been previously reported. This new understanding of oil and gas degradation by uncultured marine microbes suggested that the global ocean harbors a diversity of hydrocarbon-degrading bacteria, which can act as primary agents regulating ecosystem health.

Howe, Kathryn L.↗

“Translational Opportunities in CPS Transportation”

This talk will describe opportunities for translating research from open-road experiments with modified adaptive cruise controllers. The data and controllers used for the previous research are based on use cases and example drives from within the US. We explore processes, data sharing, experiments, and other techniques that we think will drive trans-Pacific partnerships with driving data.

Sprinkle, Jonathan↗

Two Ultra-faint Milky Way Stellar Systems Discovered in Early Data from the DECam Local Volume Exploration Survey

We report the discovery of two ultra-faint stellar systems found in early data from the DECam Local Volume Exploration survey (DELVE). The first system, Centaurus I (DELVE J1238-4054), is identified as a resolved overdensity of old and metal-poor stars with a heliocentric distance of ${\rm D}_{\odot} = 116.3_{-0.6}^{+0.6}$ kpc, a half-light radius of $r_h = 2.3_{-0.3}^{+0.4}$ arcmin, an age of $\tau > 12.85$ Gyr, a metallicity of $Z = 0.0002_{-0.0002}^{+0.0001}$, and an absolute magnitude of $M_V = -5.55_{-0.11}^{+0.11}$ mag. This characterization is consistent with the population of ultra-faint satellites, and confirmation of this system would make Centaurus I one of the brightest recently discovered ultra-faint dwarf galaxies. Centaurus I is detected in Gaia DR2 with a clear and distinct proper motion signal, confirming that it is a real association of stars distinct from the Milky Way foreground; this is further supported by the clustering of blue horizontal branch stars near the centroid of the system. The second system, DELVE 1 (DELVE J1630-0058), is identified as a resolved overdensity of stars with a heliocentric distance of ${\rm D}_{\odot} = 19.0_{-0.6}^{+0.5} kpc$, a half-light radius of $r_h = 0.97_{-0.17}^{+0.24}$ arcmin, an age of $\tau = 12.5_{-0.7}^{+1.0}$ Gyr, a metallicity of $Z = 0.0005_{-0.0001}^{+0.0002}$, and an absolute magnitude of $M_V = -0.2_{-0.6}^{+0.8}$ mag, consistent with the known population of faint halo star clusters. Here, given the low number of probable member stars at magnitudes accessible with Gaia DR2, a proper motion signal for DELVE 1 is only marginally detected. We compare the spatial position and proper motion of both Centaurus I and DELVE 1 with simulations of the accreted satellite population of the Large Magellanic Cloud (LMC) and find that neither is likely to be associated with the LMC.

79 ASTRONOMY AND ASTROPHYSICS↗

Model Choice Metrics to Optimize Profile-QSAR Performance

Predicting molecular activity against protein targets is difficult because of the paucity of experimental data. Approaches like multitask modeling and collaborative filtering seek to improve model accuracy by leveraging results from multiple targets, but are limited because different compounds are measured with different assays, leading to sparse data matrices. Profile-QSAR (pQSAR) 2.0 addresses this problem by fitting a series of partial least squares models for each target, using as features the predictions from single-task models on the remaining targets. Here, this method has been shown to produce better results than single task and multitask models. However, the factors determining the success of pQSAR 2.0 have as yet not been characterized. In this paper we examine the experimental conditions that lead to better pQSAR models. We limit the amount of data available to the method by retraining with decreasing amounts of data and explore the model’s ability to generalize to compounds that have never been assayed. Finally, we look at the properties of training data needed to demonstrate pQSAR improvement.

Biological and medical sciences, Computer science↗

Integrating Data Centers and Grid Technologies at Scale

This presentation focuses on the challenge of integrating AI-driven data centers with the power grid at scale. It examines the AI data center capacity challenge and the role of new Medium Voltage Direct Current (MVDC) and other grid-enhancing technologies in enabling efficient and reliable power delivery. The session will highlight the National Laboratory of the Rockies' ARIES capabilities and planning tools, along with collaborative examples involving Verrus, Compass, and Schneider through the Agora test bed for grid-friendly data center evaluations, and ON. Energy for UPS evaluation. It will showcase the NLR Stable Grid Platform for studying oscillations caused by large-scale data centers, along with planning tools to assess grid security and reliability. Additionally, the presentation covers reconductoring strategies to increase grid capacity and explores innovative data center architectures, including the Advanced DC Architectures with Power-electronic Transformers (ADAPT) platform, which enables testing of complete DC architectures for data centers.

24 POWER TRANSMISSION AND DISTRIBUTION↗

2020 Annual Technology Baseline (ATB) Cost and Performance Data for Transportation Technologies

The 2020 Transportation Annual Technology Baseline (ATB) provides detailed cost and performance data, estimates, and assumptions for vehicle and fuel technologies in the United States. The Transportation ATB includes current and projected estimates through 2050 for light-duty vehicle technologies as well as conventional and alternative fuels. This excel files include vehicle, fuel, and joined vehicle and fuel data with calculated levelized cost and emission values associated with the 2020 Transportation ATB. NREL has also provided a Tableau workbook to further explore the data. A website documents this data at https://atb.nlr.gov .

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Detecting Anomalous Computation with RNNs on GPU-Accelerated HPC Machines

This paper presents a workload classification framework that discriminates illicit computation from authorized workloads on GPU-accelerated HPC systems. As such systems become more and more powerful, they are exploited by attackers to run malicious and for-profit programs that typically require extremely high computing ability to be successful. Our classification framework leverages the distinctive signatures between illicit and authorized workloads, and explore machine learning methods to learn the workloads and classify them. The framework uses lightweight, non-intrusive workload profiling to collect model input data, and explores multiple machine learning methods, particularly recurrent neural network (RNN) that is suitable for online anomalous workload detection. Evaluation results on three generations of GPU machines demonstrate that the workload classification framework can tell apart the illicit authorized workloads with a high accuracy of over 95%.

Pengfei, Zou↗

Improved Earthquake Source Parameters with 3D Wavespeed Models in California and Nevada

Seismic tomography harnesses earthquake data to explore the inaccessible structure of the Earth. Adjoint waveform tomography (AWT), a method of seismic tomography, updates the tomographic model by optimizing the fit between observed earthquake data and synthetic waveforms. The synthetic data are calculated by solving the wave equation through a given 3D model. An important requirement to calculating synthetics is the source information (location, centroid time, depth, and moment tensor). Errors in source information affect the quality of the synthetics produced, which in turn can limit how structure can be inferred in the AWT workflow. Here, to test the effect of updating source information, we used MTTime (Chiang, 2020), a time-domain full-waveform moment tensor inversion code, to calculate the moment tensors and depths of 118 earthquakes that occurred in California and Nevada over a 20-yr period. We calculated 3D Green’s functions using a 3D seismic wavespeed model of California and Nevada (Doody et al., 2023b). We show that the inverted solutions provide better waveform fits than the Global Centroid Moment Tensor catalog and increase usable, well-correlated data by up to 7%. Therefore, we argue that recalculating source parameters should be considered in AWT workflows, particularly for smaller magnitude events (⁠M w > 5.0).

58 GEOSCIENCES↗

Hawai‘i Supernova Flows: a peculiar velocity survey using over a Thousand Supernovae in the near-infrared

ABSTRACT We introduce the Hawai‘i Supernova Flows project and present summary statistics of the first 1217 astronomical transients observed, 668 of which are spectroscopically classified Type Ia Supernovae (SNe Ia). Our project is designed to obtain systematics-limited distances to SNe Ia while consuming minimal dedicated observational resources. To date, we have performed almost 5000 near-infrared (NIR) observations of astronomical transients and have obtained spectra for over 200 host galaxies lacking published spectroscopic redshifts. In this survey paper, we describe the methodology used to select targets, collect/reduce data, calculate distances, and perform quality cuts. We compare our methods to those used in similar studies, finding general agreement or mild improvement. Our summary statistics include various parametrizations of dispersion in the Hubble diagrams produced using fits to several commonly used SN Ia models. We find the lowest dispersions using the SNooPy package’s EBV_model2, with a root mean square deviation of 0.165 mag and a normalized median absolute deviation of 0.123 mag. The full utility of the Hawai‘i Supernova Flows data set far exceeds the analyses presented in this paper. Our photometry will provide a valuable test bed for models of SN Ia incorporating NIR data. Differential cosmological studies comparing optical samples and combined optical and NIR samples will have increased leverage for constraining chromatic effects like dust extinction. We invite the community to explore our data by making the light curves, fits, and host galaxy redshifts publicly accessible.

Do, Aaron (ORCID:0000000334297845)↗

Data-Driven Strategies for Accelerated Materials Design

The ongoing revolution of the natural sciences by the advent of machine learning and artificial intelligence sparked significant interest in the material science community in recent years. The intrinsically high dimensionality of the space of realizable materials makes traditional approaches ineffective for large-scale explorations. Modern data science and machine learning tools developed for increasingly complicated problems are an attractive alternative. An imminent climate catastrophe calls for a clean energy transformation by overhauling current technologies within only several years of possible action available. Tackling this crisis requires the development of new materials at an unprecedented pace and scale. For example, organic photovoltaics have the potential to replace existing silicon-based materials to a large extent and open up new fields of application. In recent years, organic light-emitting diodes have emerged as state-of-the-art technology for digital screens and portable devices and are enabling new applications with flexible displays. Reticular frameworks allow the atom-precise synthesis of nanomaterials and promise to revolutionize the field by the potential to realize multifunctional nanoparticles with applications from gas storage, gas separation, and electrochemical energy storage to nanomedicine. In the recent decade, significant advances in all these fields have been facilitated by the comprehensive application of simulation and machine learning for property prediction, property optimization, and chemical space exploration enabled by considerable advances in computing power and algorithmic efficiency. In this Account, we review the most recent contributions of our group in this thriving field of machine learning for material science. We start with a summary of the most important material classes our group has been involved in, focusing on small molecules as organic electronic materials and crystalline materials. Specifically, we highlight the data-driven approaches we employed to speed up discovery and derive material design strategies. Subsequently, our focus lies on the data-driven methodologies our group has developed and employed, elaborating on high-throughput virtual screening, inverse molecular design, Bayesian optimization, and supervised learning. We discuss the general ideas, their working principles, and their use cases with examples of successful implementations in data-driven material discovery and design efforts. Furthermore, we elaborate on potential pitfalls and remaining challenges of these methods. Finally, we provide a brief outlook for the field as we foresee increasing adaptation and implementation of large scale data-driven approaches in material discovery and design campaigns.

36 MATERIALS SCIENCE↗

Phase Selection Rules of Multi‐Principal Element Alloys

Abstract Computational prediction of phase stability of multi‐principal element alloys (MPEAs) holds a lot of promise for rapid exploration of the enormous design space and autonomous discovery of superior structural and functional properties. Regardless of many plausible works that rely on phenomenological theory and machine learning, precise prediction is still limited by insufficient data and the lack of interpretability of some machine learning algorithms, e.g., convolutional neural network. In this work, a comprehensive approach is presented, encompassing the development of a complete dataset that contains 72 387 density functional theory calculations, as well as a predictive global phenomenological descriptor. The phase selection descriptor, based on atomic electronegativity and valence electron concentration, significantly outperforms the widely used valence electron concentration, excelling in both accuracy (with an f1 score of 63% compared to 47%) and its ability to predict the HCP phase (0.48 recall compared to 0). The comprehensive data mining on the global design space of 61 425 quaternary MPEAs made from 28 possible metals, together with the phenomenological theory and physical interpretation, will set up a solid computational science foundation for data‐driven exploration of MPEAs.

Chemistry↗

On-Line Monitoring of Gas-Phase Molecular Iodine Using Raman and Fluorescence Spectroscopy Paired with Chemometric Analysis

Molten salt reactors (MSRs) have the potential to safely support green energy goals. However, licensing and deployment of these systems will be aided through development of new technology. This includes on-line monitoring tools for real-time compositional analysis. Of particular interest is quantifying iodine within reactor off-gas streams to support design and operational control of reactor off-gas treatment systems. Here we discuss the development of advanced Raman spectroscopy systems for the on-line analysis of I 2(g) within the gas phase. Signal response is explored with two Raman instruments utilizing a 532 nm and a 671 nm excitation source, as a function of I 2(g) pressure and temperature. Furthermore, the applicability of chemometric modeling for advanced analysis of data is explored. Raman spectroscopy paired with chemometric analysis is demonstrated to be a powerful route to analyzing I 2(g) composition within the gas phase, which lays the foundation for applications within molten salt reactor off-gas analysis and other significant chemical processes producing iodine species.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Cosmological preference for a negative neutrino mass

The most precise determination of the sum of neutrino masses from cosmological data, derived from analysis of the cosmic microwave background (CMB) and baryon acoustic acoustic oscillations (BAO) from the Dark Energy Spectroscopic Instrument (DESI), favors a value below the minimum inferred from neutrino flavor oscillation experiments. We explore which data is most responsible of this puzzling aspect of the current constraints on neutrino mass and whether it is related to other anomalies in cosmology. We demonstrate conclusively that the preference for negative neutrino masses is a consequence of larger than expected lensing of the CMB in both the two- and four-point lensing statistics. Furthermore, we show that this preference is robust to changes in likelihoods of the BAO and CMB optical depth analyses given the available data. We then show that this excess clustering is not easily explained by changes to the expansion history and is likely distinct from the preference for for dynamical dark energy in DESI BAO data. Finally, we discuss how future data may impact these results, including an analysis of Planck CMB with mock DESI 5-year data. Here, we conclude that the negative neutrino mass preference is likely to persist even as more cosmological data is collected in the near future.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Patience is a virtue: A data-driven analysis of rooftop solar PV permitting timelines in the United States

© 2020 The Authors Local permitting can ensure the safe installation and operation of rooftop solar photovoltaic (PV) systems. At the same time, burdensome local permitting processes and local variation in requirements may pose challenges to PV deployment. In this article, we explore new data on the durations between key steps in the PV permitting process in the United States. The data suggest that a typical customer can expect to wait around 25–100 days from permit application until an installed system passes inspection. Permit durations vary significantly across jurisdictions, due in part to differences in local permitting policies. However, permit durations vary as significantly within jurisdictions as across them, in part due to significant variation across installers, suggesting that installer strategies and practices play an important role in permitting timelines. Permit durations have declined over time, reflecting progress from permit streamlining policies and jurisdiction learning-by-doing, though durations have stabilized in recent years. The data suggest that typical PV customers still face long and uncertain permitting timelines in the United States.

O'Shaughnessy, E↗

Assessments of data centers for provision of frequency regulation

There are numerous opportunities for data centers to participate in demand response programs considering their large energy capacities, flexible working environments and workloads, redundant design and operation, etc. As a type of demand response, frequency regulation requires fast response, and its potential is not fully explored by data centers yet. This paper proposes a synergistic control strategy for data center frequency regulation which uses both IT and cooling systems. It combines power management techniques at the server level with control of the chilled water supply temperature to track the regulation signal from the electrical market. A frequency regulation flexibility factor is also proposed to increase the IT capacity for frequency regulation. The performance of the control strategy is studied through numerical simulations using an equation based object-oriented Modelica platform designed for data centers. Simulation results show that with well-tuned control parameters, data centers can provide frequency regulation service in both regulation up and down. The performance of data centers in providing frequency regulation service is largely influenced by the regulation capacity bid, frequency regulation flexibility factor, workload condition, and cooling mode of the cooling system, and not significantly influenced by the time constant of chillers. In addition, compared with a server-only control strategy, the proposed synergistic control strategy can provide an extra regulation capacity of 3% of the design power when chillers are activated. Here, when chillers are deactivated, both strategies have a similar regulation capacity.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

SAS4A/SASSYS-1 Validation with EBR-II Tests Performed During the SHRT Testing Program (Rev. 2)

The EBR-II sodium-cooled fast reactor supported a wide range of fast reactor research activities. The final fifteen years of operations at EBR-II were used for experiments and tests to demonstrate the importance of passive safety in liquid metal reactors and the capability of such systems to provide strong passive responses during off-normal conditions. The Shutdown Heat Removal Test program at EBR-II provided test data supporting the validation of computer codes for design, licensing, and operation of LMRs, among other objectives. The program included nearly sixty tests, most of which were protected and unprotected loss of flow and loss of heat sink tests run at various initial powers and flow rates. This report documents the results of modeling and simulation of tests from the Shutdown Heat Removal Test program to support validation of Argonne’s SAS4A/SASSYS-1 fast reactor safety analysis code. SHRT-17 and SHRT-6 were protected loss of flow tests where a loss of electrical power to all sodium coolant pumps was simulated to demonstrate the effectiveness of natural circulation cooling characteristics. SHRT-45R, SHRT-39, and SHRT-43R were unprotected loss of flow tests where the control rod scram function of the plant protection system was disabled to demonstrate the effectiveness of EBR-II’s passive reactivity feedbacks. BOP-301 and BOP-302R were unprotected loss of heat sink tests that demonstrated how reactivity feedback effects driven by changes to the core inlet temperature can shut down the fission process. PICT-5/6 demonstrated how core power could be increased and decreased via changes in the intermediate sodium and feedwater flow rates. This validation activity leverages modeling and simulation efforts that originated under an International Atomic Energy Agency Coordinated Research Project for benchmark analysis of the EBR-II Shutdown Heat Removal Tests. The SAS4A/SASSYS-1 models originally developed for the CRP have been modernized to utilize new code features and modeling practices. Differences between the predicted and measured test data were explored and the causes of these differences were identified as being due to modeling approximations, uncertainty in transient component behavior, and instrumentation uncertainties. Overall, it was concluded that for all tests simulated there was good agreement between the flow, power, and temperature predictions and the measured data.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗