Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Transferable predictions of energetic and structural properties for refractory solid solution alloys across chemical compositions

We present a data-efficient approach to train graph neural networks (GNNs) on density functional theory (DFT) data for accurate and transferable predictions of energetic and structural properties of refractory solid solution alloys in the niobium-tantalum-vanadium (Nb-Ta-V) chemical space. We start by training the GNN model only on DFT data that describes refractory binary alloys niobium-tantalum (Nb-Ta), niobium-vanadium (Nb-V), and tantalum-vanadium (Ta-V) to predict formation enthalpy and root mean squared displacement. Once trained, the GNN predictions are tested on DFT data describing refractory ternary alloys Nb-Ta-V. While, unsurprisingly, direct transferability from binary to ternary is not sufficiently accurate, augmenting the training with only 1% of the available ternary data (uniformly distributed across the entire range of chemical compositions) improves significantly the quality of the GNN predictions. For comparison, we assess the transferability in the opposite direction by training GNN models on ternary Nb-Ta-V data and making predictions on binaries Nb-Ta, Nb-V, and Ta-V, which exhibits notably higher predictive errors. The proposed methodology, which favors transferability from lower-component to higher-component alloys, offers an efficient path towards avoiding the curse of dimensionality incurred when collecting DFT data for discovery and design of multi-component disordered alloys.

Density functional theory calculations↗

High-throughput computation of electric polarization in solids via Berry flux diagonalization

Electric polarization in the absence of an externally applied electric field is a key property of polar materials, but the standard interpolation-based ab initio approach to compute polarization differences within the modern theory of polarization presents challenges for automated high-throughput calculations. Berry flux diagonalization [J. Bonini et al., Phys. Rev. B 102, 045141 (2020)] has been proposed as an efficient and reliable alternative, though it has yet to be widely deployed. Here, we assess Berry flux diagonalization using ab initio calculations of a large set of materials, introducing and validating heuristics that ensure branch alignment with a minimal number of intermediate interpolated structures. Our automated implementation of Berry flux diagonalization succeeds in cases where prior interpolation-based workflows fail due to band-gap closures or branch ambiguities. Benchmarking with ab initio calculations of 176 candidate ferroelectrics, we demonstrate the efficacy of the approach on a broad range of insulating materials and obtain accurate effective polarization values with fewer interpolated structures than prior automated interpolation-based workflows. Our real-space heuristics that can predict gauge stability a priori from ionic displacements enable a general automated framework for reliable polarization calculations and efficient high-throughput screening of chemically and structurally diverse polar insulators. These results establish Berry flux diagonalization as a robust and efficient method to compute the effective polarization of solids and to accelerate the data-driven discovery of functional polar materials.

Poteshman, Abigail N. [University of Chicago, IL (↗

Superionic conduction in solid polymer electrolytes – decoupling ion transport from segmental relaxation

Solvent-free, solid polymer electrolytes (SPEs) are promising candidates for next-generation, electrochemical energy storage systems due to their potential to enhance safety and performance, enable flexible device architectures, and streamline manufacturing processes. Conventional SPEs suffer from limited ionic conductivity due to the strong coupling between ion transport and (generally slow) polymer segmental relaxation. The realization of superionic conduction in SPEs, in which ions move faster than the structural relaxation of the polymers, requires a shift in design principles to promote this type of decoupled ion motion. In this perspective, we discuss how polymer architecture, ion–ion correlations, and ion–polymer interactions can unlock superionic behavior. We highlight several key design features, such as crystallinity, bulky side groups, high molecular weight, and percolating ionic aggregation, with a focus on creating low-barrier transport pathways in various polymer systems. We also demonstrate opportunities to combine polymer chemistry and data science through high-throughput and automated screening approaches to reveal how phase behavior, ion dynamics, and ionic interactions govern transport, thereby potentially enabling data-driven discovery of superionic polymer electrolyte materials.

Yang, Mengying [Univ. of Delaware, Newark, DE (Uni↗

Data-driven multi-element substitution of TiFe alloys for tunable thermodynamics and enhanced activation behaviour for hydrogen storage

Due to their high volumetric hydrogen storage capacity under moderate storage conditions, TiFe alloys have been widely investigated as candidates for practical solid-state hydrogen storage. Partially substituting Ti or Fe sites can improve the key characteristics of TiFe alloys, such as the first hydrogen absorption step (activation) and the equilibrium hydrogen pressure (thermodynamic properties). However, the selection of substitution elements has heavily relied on intuition and trial-and-error. Also, conventional substitution strategies have mainly focused on single-element substitution within the TiFe alloy, limiting the design space and tunability for target applications. Here, to address this limitation, we report a multi-element substitution strategy motivated by an efficient, data-driven machine learning (ML) approach combined with corroborating density functional theory (DFT) calculations. Our models successfully predict experimentally measured hydride stability in five selected alloys using only compositional descriptors. Most importantly, the multi-element substitution leads to enhanced activation properties compared to pure TiFe, achieving near room-temperature activation behaviour. This work provides a method for on-demand tuning of hydrogen storage and activation properties, which may have broad implications for data-driven discovery of energy storage materials.

Cho, YongJun [Korea Advanced Institute Science and↗

Methods for Causal Discovery

SAND2025-11742O Methods for Causal Discovery is a software tool that is used for causal discovery from data, including predicting and visualizing directed acyclic graphs from data using traditional machine learning techniques. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

SciDAC↗

Quasars Acting as Strong Lenses Found in DESI DR1

Quasars acting as strong gravitational lenses offer a rare opportunity to probe the redshift evolution of scaling relations between supermassive black holes and their host galaxies, particularly the M$_{BH}$–M$_{host}$ relation. Using these powerful probes, the mass of the host galaxy can be precisely inferred from the Einstein radius θ$_{E}$. Using 812,118 quasars from DESI DR1 (0.03 ≤ z ≤ 1.8), we searched for quasars lensing higher-redshift galaxies by identifying background emission-line features in their spectra. To detect these rare systems, we trained a convolutional neural network (CNN) on mock lenses constructed from real DESI spectra of quasars and emission-line galaxies (ELGs), achieving a high classification performance (AUC = 0.99). We also trained a regression network to estimate the redshift of the background ELG. Applying this pipeline, we identified seven high-quality (Grade A) lens candidates, each exhibiting a strong [O II] doublet at a higher redshift than the foreground quasar; four candidates additionally show Hβ, [O III] λ4959, and [O III] λ5007 emission. These results significantly expand the sample of quasar lens candidates beyond the 12 identified and 3 confirmed in previous work and demonstrate the potential for scalable, data-driven discovery of quasars as strong lenses in upcoming spectroscopic surveys.

McArthur, Everett [Stanford U., Phys. Dept.; KIPAC↗

Enabling Early Transient Discovery in LSST via Difference Imaging with DECam

We present SLIDE, a pipeline that enables transient discovery in data from the Vera C. Rubin Observatory’s Legacy Survey of Space and Time (LSST), using archival images from the Dark Energy Camera as templates for difference imaging. We apply this pipeline to the recently released Data Preview 1 (DP1; the first public release of Rubin commissioning data) and search for transients in the resulting difference images. The image subtraction, photometry extraction, and transient detection are all performed on the Rubin Science Platform. We demonstrate that SLIDE effectively extracts clean photometry by circumventing poor or missing LSST templates. We identified 29 previously unreported transients, 12 of which would not have been detected based on the DP1 DiaObject catalog. SLIDE will be especially useful for transient analysis in the early years of LSST, when template coverage will be largely incomplete or when templates may be contaminated by transients present at the time of acquisition. We present multiband light curves for a sample of known transients, along with new transient candidates identified through our search. Finally, we discuss the prospects of applying this pipeline during the main LSST survey. Our pipeline is broadly applicable and will support studies of all transients with slowly evolving phases.

Dong, Yize 一泽董 [Harvard-Smithsonian Center for Ast↗

YeastWGD2025

Supplementary data for Discovery of additional ancient genome duplications in yeasts wgd_syn / - directory containing wgd syn output for all contiguous genomes [dataset] Tree - phylogeny [dataset]Duplications - duplication table from OrthoFinder output KOannotations - KEGG annotations used for enrichment analysis IPRannotations - InterPro annotations used for enrichment analysis DipodascalesOrthogroups - formatted orthogroup assignments for Dipodascales genes.fa and .gff3 files for each new genome assembly are also provided, those these are not required to replicate the analysis

Genomics↗

Understanding and Estimating Error Propagation in Neural Networks for Scientific Data Analysis

Neural networks are increasingly integrated into scientific discovery, where input data reduction and model quantization play a key role in accelerating inference. However, understanding and mitigating the impact of these techniques on output error is critical for ensuring reliable results, particularly in tasks demanding high numerical precision. This paper introduces a comprehensive framework for optimizing neural network inference in scientific computing by combining data reduction and weight quantization while maintaining error-controlled outcomes. We develop theoretical analyses to bound error propagation under these reductions and propose a framework that balances computational performance with error constraints. Evaluation on real-world learning-based combustion simulations and satellite image classification demonstrates that our derived error bounds accurately predict observed errors while enabling significant computational speedup under our framework. This work highlights the potential for further leveraging advancements in modern lossy compression algorithms and hardware accelerators that support lower-precision formats.

He, Weiming [New Jersey Institute of Technology]↗

Leveraging data mining, active learning, and domain adaptation for efficient discovery of advanced oxygen evolution electrocatalysts

Developing advanced catalysts for acidic oxygen evolution reaction (OER) is crucial for sustainable hydrogen production. This study presents a multistage machine learning (ML) approach to streamline the discovery and optimization of complex multimetallic catalysts. Our method integrates data mining, active learning, and domain adaptation throughout the materials discovery process. Unlike traditional trial-and-error methods, this approach systematically narrows the exploration space using domain knowledge with minimized reliance on subjective intuition. Then, the active learning module efficiently refines element composition and synthesis conditions through iterative experimental feedback. The process culminated in the discovery of a promising Ru-Mn-Ca-Pr oxide catalyst. Our workflow also enhances theoretical simulations with domain adaptation strategy, providing deeper mechanistic insights aligned with experimental findings. By leveraging diverse data sources and multiple ML strategies, we demonstrate an efficient pathway for electrocatalyst discovery and optimization. This comprehensive, data-driven approach represents a paradigm shift and potentially benchmark in electrocatalysts research.

Science & Technology - Other Topics↗

Machine Learning on Heterogeneous, Edge, and Quantum Hardware for Particle Physics (ML-HEQUPP)

The next generation of particle physics experiments will face a new era of challenges in data acquisition, due to unprecedented data rates and volumes along with extreme environments and operational constraints. Harnessing this data for scientific discovery demands real-time inference and decision-making, intelligent data reduction, and efficient processing architectures beyond current capabilities. Crucial to the success of this experimental paradigm are several emerging technologies, such as artificial intelligence and machine learning (AI/ML) and silicon microelectronics, and the advent of quantum algorithms and processing. Their intersection includes areas of research such as low-power and low-latency devices for edge computing, heterogeneous accelerator systems, reconfigurable hardware, novel codesign and synthesis strategies, readout for cryogenic or high-radiation environments, and analog computing. This white paper presents a community-driven vision to identify and prioritize research and development opportunities in hardware-based ML systems and corresponding physics applications, contributing towards a successful transition to the new data frontier of fundamental science.

Gonski, Julia [SLAC]↗

Toward the Neutrino Discovery Platform: An Auditable, Uncertainty-Bearing Toolchain for MINERvA Open-Data Cross-Section Analysis

The Neutrino Discovery Platform (NDP) aims to accelerate DUNE-era science by making the neutrino program's existing datasets analyzable through fast, reproducible, and auditable workflows. We report a working version of two of its layers, data curation and agentic orchestration, built and tested end to end on MINERvA open data. The guiding lesson throughout is that a cross section is a measurement, and not just a plotted shape, only if it carries a defensible systematic-uncertainty budget, a trustworthy unfolding, and a reproducible record. Using a single medium-energy playlist pair from the MINERvA open-data release (about $2.05\times10^{17}$ protons on target of data), we first reproduced the shapes of two published charged-current inclusive $\nu_\mu$ measurements through a complete extraction ladder: selection, background subtraction, D'Agostini unfolding, efficiency correction, and flux normalization. These shape-level reproductions ran and tracked the published results, but they lacked the systematic-uncertainty machinery that defines a MINERvA cross section. To supply it, we vendored and built the MINERvA Analysis Toolkit and developed a many-universe systematic-uncertainty tool that produces a portable covariance artifact, a parallel event-loop runner, and a per-run auditability harness. Validated against a published covariance release, the toolchain reproduces the released statistical, flux, and muon-energy-scale terms and shows that they account for roughly 63\% of the total variance, with the remainder unreleased. Using this same infrastructure, we then performed a measurement of our own design, the hadronic recoil-energy distribution of low-energy ($E_\nu<2.5$~GeV) charged-current inclusive events, and found data/simulation shape agreement of $\chi^2/\mathrm{ndf}=1.26$. Together these results show that the platform supports original physics and not only reproductions.

Breaux, Auto [Tulane U. (main)]↗

Scalable Hybrid Learning Techniques for Scientific Data Compression

Data compression is becoming critical for storing scientific data because many scientific applications need to store large amounts of data and post process this data for scientific discovery. Unlike image and video compression algorithms that limit errors to primary data (PD), scientists require compression techniques that accurately preserve derived quantities of interest (QoIs). Here, this article presents a physics-informed compression technique implemented as an end-to-end, scalable, GPU-based pipeline for data compression that addresses this requirement. Our hybrid compression technique combines machine learning techniques and standard compression methods. Specifically, we combine an autoencoder, an error-bounded lossy compressor to provide guarantees on raw data error, and a constraint satisfaction post-processing step to preserve the QoIs within a minimal error (generally less than floating point error). The effectiveness of the data compression pipeline is demonstrated by compressing nuclear fusion simulation data generated by a large-scale fusion code, XGC, which produces hundreds of terabytes of data in a single day. Our approach works within the ADIOS framework and results in compression by a factor of more than 150 while requiring only a few percent of the computational resources necessary for generating the data, making the overall approach highly effective for practical scenarios.

ITER↗

G2PDeep-v2: A Web-Based Deep-Learning Framework for Phenotype Prediction and Biomarker Discovery for All Organisms Using Multi-Omics Data

Multi-omics data offers rich insights into complex traits across organisms, yet integrating and analyzing these datasets for phenotype prediction and marker discovery remains challenging. Researchers need accessible tools that combine deep learning, hyperparameter optimization, visualization, and downstream analysis in a unified web platform. To address this, we developed G2PDeep-v2, a web-based platform powered by deep learning for phenotype prediction and marker discovery from multi-omics data across a wide range of organisms, including humans and plants. The server provides multiple services for researchers to create deep-learning models through an interactive interface and train these models using an automated hyperparameter tuning algorithm on high-performance computing resources. Users can visualize the results of phenotype and markers predictions and perform Gene Set Enrichment Analysis for the significant markers to provide insights into the molecular mechanisms underlying complex diseases, conditions and other biological phenotypes being studied.

59 BASIC BIOLOGICAL SCIENCES↗

All-sky Neutrino Point-source Search with IceCube Combined Track and Cascade Data

Despite extensive efforts, discovery of high-energy astrophysical neutrino sources remains elusive. We present an event-level simultaneous maximum likelihood analysis of tracks and cascades using IceCube data collected from 2008 April 6 to 2022 May 23 to search the whole sky for neutrino sources, and using a source catalog, for coincidence of neutrino emission with gamma-ray emission. This is the first time a simultaneous fit of different detection channels is used to conduct a time-integrated all-sky scan with IceCube. Combining all-sky tracks, with superior pointing power and sensitivity in the northern sky, with all-sky cascades, with good energy resolution and sensitivity in the southern sky, we have developed the most sensitive point-source search to date by IceCube that targets the entire sky. The most significant point in the northern sky aligns with NGC 1068, a Seyfert II galaxy, which, from the catalog search, shows a 3.5σ excess over background after accounting for trials. The most significant point in the southern sky does not align with any source in the catalog and is not significant after accounting for trials. A search for the single most significant Gaussian flare at the locations of NGC 1068, PKS 1424+240, and the southern highest-significance point shows results consistent with expectations for steady emission. Notably, this is the first time that a flare shorter than four years has been excluded as being responsible for NGC 1068’s emergence as a neutrino source. Our results show that combining tracks and cascades when conducting neutrino source searches improves sensitivity and can lead to new discoveries.

Abbasi, R. [Loyola University, Chicago, IL (United↗

Privacy Preservation from High-Performance Computing to Autonomous Science [Industrial and Governmental Activities]

High-Performance Computing (HPC) and Leadership-Class Supercomputing are driving forces behind scientific advancements, enabling researchers to tackle complex challenges in physics, chemistry, biology, and engineering. These systems power vast simulations and data analyses, fueling discoveries in fields ranging from materials science to climate modeling. However, their use often involves processing sensitive data—such as proprietary industry simulations, biomedical records, and national security computations—posing significant privacy concerns. In conclusion, this issue is amplified in collaborative environments like Department of Energy (DOE) user facilities, where HPC resources are shared across institutions to foster innovation.

Kotevska, Olivera [Oak Ridge National Laboratory (↗

Pre-training Vision Models for the Classification of Alerts from Wide-field Time-domain Surveys

Modern wide-field time-domain surveys facilitate the study of transient, variable and moving phenomena by conducting image differencing and relaying alerts to their communities. Machine learning tools have been used on data from these surveys and their precursors for more than a decade, and convolutional neural networks (CNNs), which make predictions directly from input images, saw particularly broad adoption through the 2010s. Since then, continually rapid advances in computer vision have transformed the standard practices around using such models. It is now commonplace to use standardized architectures pre-trained on large corpora of everyday images (e.g., ImageNet). In contrast, time-domain astronomy studies still typically design custom CNN architectures and train them from scratch. Here, we explore the effects of adopting various pre-training regimens and standardized model architectures on the performance of alert classification. We find that the resulting models match or outperform a custom, specialized CNN like what is typically used for filtering alerts. Moreover, our results show that pre-training on galaxy images from Galaxy Zoo tends to yield better performance than pre-training on ImageNet or training from scratch. We observe that the design of standardized architectures are much better optimized than the custom CNN baseline, requiring significantly less time and memory for inference despite having more trainable parameters. On the eve of the Legacy Survey of Space and Time and other image-differencing surveys, these findings advocate for a paradigm shift in the creation of vision models for alerts, demonstrating that greater performance and efficiency, in time and in data, can be achieved by adopting the latest practices from the computer vision field.

79 ASTRONOMY AND ASTROPHYSICS↗

A Pride of Satellites in the Constellation Leo? Discovery of the Leo VI Milky Way Satellite Ultra-faint Dwarf Galaxy with DELVE Early Data Release 3

Abstract We report the discovery and spectroscopic confirmation of an ultra-faint Milky Way satellite in the constellation of Leo. This system was discovered as a spatial overdensity of resolved stars observed with Dark Energy Camera (DECam) data from an early version of the third data release of the DECam Local Volume Exploration (or DELVE) survey. The low luminosity ( M V = − 3.5 6 − 0.37 + 0.47 ; L V = 230 0 − 700 + 1200 L ⊙ ), large size ( R 1 / 2 = 9 0 − 30 + 30 pc), and large heliocentric distance ( D = 11 1 − 6 + 9 kpc) are all consistent with the population of ultra-faint dwarf galaxies (UFDs). Using Keck/DEIMOS observations of the system, we were able to spectroscopically confirm nine member stars, while measuring a tentative mass-to-light ratio of 70 0 − 500 + 1400 M ⊙ / L ⊙ and a nonzero metallicity dispersion of σ [ Fe / H ] = 0.1 9 − 0.11 + 0.14 , further confirming Leo VI’s identity as a UFD. While the system has a highly elliptical shape, ϵ = 0.5 4 − 0.29 + 0.19 , we do not find any conclusive evidence that it is tidally disrupting. Moreover, despite the apparent on-sky proximity of Leo VI to members of the proposed Crater-Leo infall group, its smaller heliocentric distance and inconsistent position in energy–angular momentum space make it unlikely that Leo VI is part of the proposed infall group.

79 ASTRONOMY AND ASTROPHYSICS↗