Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “large-scale data association”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Risk-Aware Measurement Synchronization and Recovery for DSSE With Heterogeneous Data Sources

Power distribution systems are increasingly integrating heterogeneous sensors with varying data reporting rates and types, which pose challenges to achieving observability at the desired temporal resolution of distribution system state estimation (DSSE). Multisensor failures caused by extreme events exacerbate these issues, introducing substantial uncertainties into DSSE. This article proposes a novel solution to these challenges by ensuring high-resolution system observability despite heterogeneous data sources and multisensor failures. First, a deep learning architecture combining long short-term memory (LSTM) and graph convolutional network (GCN) is employed to synchronize meters with different reporting rates, aiming to achieve system observability. A random-walk-model-based approach is introduced to generate pseudo-measurements while properly characterizing their uncertainties under multisensor failures. Finally, a disaster-risk-informed observability metric (RiOM) is defined to quantify the uncertainty associated with state estimation results. The proposed framework offers deeper insights into the system observability on the fly compared with conventional analysis. The effectiveness of the framework is demonstrated on an IEEE standard test case and a large-scale real-world distribution feeder in mid-Minnesota in the U.S.

97 MATHEMATICS AND COMPUTING↗

Similarity Metric for Data Optimization and Efficient Training of Reactive Machine Learning Force Fields for Hydrocarbon Radiolysis

Radiolysis is a common approach to sterilize polymers, chemically modify them for upcycling, and accelerate their decomposition for recycling purposes. Reactive molecular dynamics (MD) simulations provide a powerful tool to generate atomic-level trajectories of the reactive processes and quantify radiolytic chemical degradation pathways. For this, machine learning (ML) surrogate models for reactive force fields with quantum mechanical accuracy are now widely used, which require ML training data sets that can provide information on atomic environments for target chemical systems. However, radiolysis chemistry can be highly complex and diverse, which poses significant challenges for generating training data to parametrize ML models. In this regard, we developed a method for optimizing the training data set using a cosine similarity metric to help guide training set selection for radiolysis of polyethylene, a model hydrocarbon polymer, as well as to enhance the transferability of our reactive ML force field (MLFF) to a variety of molecular and polymeric systems. Our approach performs atom-by-atom comparisons between local atomic environments to pinpoint important data points associated with rare and localized events, such as radiolysis damage within structures. We apply this approach to train the Chebyshev Interaction Model for Efficient Simulation (ChIMES) MLFF model, which expresses the atomic interaction potentials in terms of linear combinations of many-body Chebyshev polynomials. We first show that our method can reduce our training set size by ∼70% while improving overall accuracy compared to more standard MD model fitting approaches. We then validate our optimum model against diverse hydrocarbon simulation data, including simple alkanes and systems with unsaturated carbon bonds, over a wide range of thermodynamic conditions. Finally, we use our ChIMES model to perform MD simulations of radiolytic damage with large-scale systems that help avoid system size effects. Overall, our approach yields an MD force field that retains most of the accuracy of the underlying quantum method while yielding many orders of improvement in computational efficiency. In conclusion, our efforts will have impact on future hydrocarbon polymer radiolysis studies, where the chemical details of the polymer–radiation interactions can have a strong effect on the resulting products observed in experiments.

Hydrocarbons↗

Taylor approximation variance reduction for approximation errors in PDE-constrained Bayesian inverse problems

In numerous applications, surrogate models are used as a replacement for accurate parameter-to-observable mappings when solving large-scale inverse problems governed by partial differential equations (PDEs). The surrogate model may be a computationally cheaper alternative to the accurate parameter-to-observable mappings and/or may ignore additional unknowns or sources of uncertainty. The Bayesian approximation error (BAE) approach provides a means to account for the induced uncertainties and approximation errors, i.e. the errors between the accurate parameter-to-observable mapping and the surrogate. The statistics of these errors are, however, in general unknown a priori, and are thus calculated using Monte Carlo sampling. Although the sampling is typically carried out offline, i.e. before considering the data, the process can still represent a computational bottleneck. In this work, we develop a scalable computational approach for reducing the costs associated with the sampling stage of the BAE approach. Specifically, we consider the Taylor expansion of the accurate and surrogate forward models with respect to the uncertain parameter fields either as a control variate for variance reduction or as a means to directly and efficiently approximate the mean and covariance of the approximation errors. We propose efficient methods for evaluating the expressions for the mean and covariance of the Taylor approximations based on linear(-ized) PDE solves. Furthermore, the proposed approach is independent of the dimension of the uncertain parameter, depending instead on the intrinsic dimension of the data, ensuring scalability to high-dimensional problems. The potential benefits of the proposed approach are demonstrated for two high-dimensional inverse problems governed by PDE examples, namely for the estimation of a distributed Robin boundary coefficient in a linear diffusion problem, and for a coefficient estimation problem governed by a nonlinear diffusion problem.

Bayesian approximation error↗

A Parameter-masked Mock Data Challenge for Beyond-two-point Galaxy Clustering Statistics

The past few years have seen the emergence of a wide array of novel techniques for analyzing high-precision data from upcoming galaxy surveys, which aim to extend the statistical analysis of galaxy clustering data beyond the linear regime and the canonical two-point (2pt) statistics. We test and benchmark some of these new techniques in a community data challenge named “Beyond-2pt,” initiated during the Aspen 2022 Summer Program “Large-Scale Structure Cosmology beyond 2-Point Statistics,” whose first round of results we present here. The challenge data set consists of high-precision mock galaxy catalogs for clustering in real space, in redshift space, and on a light cone. Participants in the challenge have developed end-to-end pipelines to analyze mock catalogs and extract unknown (“masked”) cosmological parameters of the underlying ΛCDM models with their methods. The methods represented are density-split clustering, nearest neighbor statistics, BACCO power spectrum emulator, void statistics, LEFTfield field-level inference using effective field theory (EFT), and joint power spectrum and bispectrum analyses using both EFT and simulation-based inference. In this work, we review the results of the challenge, focusing on problems solved, lessons learned, and future research needed to perfect the emerging beyond-2pt approaches. The unbiased parameter recovery demonstrated in this challenge by multiple statistics and the associated modeling and inference frameworks supports the credibility of cosmology constraints from these methods. The challenge data set is publicly available, and we welcome future submissions from methods that are not yet represented.

Krause, Elisabeth [Univ. of Arizona, Tucson, AZ (U↗

ReMU: regional minimal updating for model-based derivative-free optimization

Derivative-free optimization (DFO) problems are optimization problems where derivative information is unavailable or extremely difficult to obtain. Model-based DFO solvers have been applied extensively in scientific computing. Powell's NEWUOA (2004) [Powell, The NEWUOA software for unconstrained optimization without derivatives, in Large-Scale Nonlinear Optimization, Nonconvex Optimization and its Applications Vol. 83, G. Di Pillo and M. Roma, eds., Springer, 2006, pp. 255–297] and Wild's POUNDerS (2014) [Wild, Solving derivative-free nonlinear least squares problems with POUNDERS, in Advances and Trends in Optimization with Engineering Applications, T. Terlaky, M.F. Anjos, and S. Ahmed, eds., SIAM, 2017, pp. 529–540] explore the numerical power of the minimal norm Hessian (MNH) model for DFO and contributed to the open discussion on building better models with fewer data to achieve faster numerical convergence. Another decade later, we propose the regional minimal updating (ReMU) models, and extend the previous models into a broader class, including the H 2 norm models [Xie and Yuan, Least H 2 norm updating of quadratic interpolation models for derivative-free trust-region algorithms, IMA J. Numer. Anal. 46 (2025), pp. 21–50]. This paper shows motivation behind ReMU models, computational details, theoretical and numerical results on particular extreme points and the barycentre of ReMU's weight coefficient region, and the associated KKT matrix error and distance. Novel metrics, such as the truncated Newton step error, are proposed to numerically understand the new models' properties. A new algorithmic strategy, based on iteratively adjusting the ReMU model type, is also proposed, and shows numerical advantages by combining and switching between the barycentric model and the classic least Frobenius norm model in an online fashion.

derivative-free trust-region methods↗

Selection function of clusters in Dark Energy Survey year 3 data from cross-matching with South Pole Telescope detections

Context. Galaxy clusters selected based on overdensities of galaxies in photometric surveys provide the largest cluster samples. However, modeling the selection function of such samples is complicated by noncluster members projected along the line of sight (projection effects) and the potential detection of unvirialized objects (contamination). Aims. We empirically constrained the magnitude of these effects by cross-matching galaxy clusters selected in the Dark Energy Survey data with the redMaPPer algorithm with significant detections in three South Pole Telescope surveys (SZ, pol-ECS, pol-500d). Methods. For matched clusters, we augmented the redMaPPer catalog with the SPT detection significance. For unmatched objects we used the SPT detection threshold as an upper limit on the SZe signature. Using a Bayesian population model applied to the collected multiwavelength data, we explored various physically motivated models to describe the relationship between observed richness and halo mass. Results. Our analysis reveals a clear preference for models with an additional skewed scatter component associated with projection effects over a purely log-normal scatter model. We rule out significant contamination by unvirialized objects at the high-richness end of the sample. While dedicated simulations offer a well-fitting calibration of projection effects, our findings suggest the presence of redshift-dependent trends that these simulations may not have captured. Our findings highlight that modeling the selection function of optically detected clusters remains a complicated challenge that requires a combination of simulation and data-driven approaches.

79 ASTRONOMY AND ASTROPHYSICS↗

Systems Analysis of Biomass and Coal Co-firing Power Plants with Deep Carbon Capture Toward Net-zero Emissions

Achieving a net-zero emission economy in the United States requires integrating diverse low-carbon and negative-emission technologies into the existing fossil fuel-dominant power fleet. Potential technologies from the low-carbon portfolio include renewable power, fossil power with carbon capture and storage (CCS), bioenergy with CCS (BECCS), and direct air capture (DAC). Renewable power is a clean energy source but has to pair with costly battery storage to provide dispatchable electricity. Fossil power with CCS offers dispatchable electricity yet still relies on DAC to offset residual emissions, even when deploying deep CCS with more than 90% CO2 capture. Coal-biomass co-firing with CCS, a subset of BECCS, is a reliable energy production technology that can be retrofitted from existing electricity generation units (EGUs). Power plant retrofit maximizes the use of the current U.S. coal power fleet without the need for large-scale deployment of new renewable power, battery storage, or DAC. Retrofitting coal-biomass co-firing with deep CCS in EGUs is a promising option, but not a universal solution. Biomass co-firing at a power plant introduces economic challenges and indirectly poses pressure on land and water resources. Meanwhile, retrofitting deep CCS affects plant efficiency and raises electricity generation costs. Overall, the technical feasibility and economic viability of plant retrofits vary across EGUs, as they are contingent upon the regional availability of biomass, unit-specific characteristics, site-specific fuel supply costs, and adjacent CO2 storage potential. Government incentives like 45Q can improve the retrofit viability, though the impact requires further quantification. A comprehensive analysis at the unit level is essential to address the question regarding the fate of the U.S. coal-fired electricity generation fleet toward the net-zero emission goal. This study conducts a systematic techno-economic-environmental assessment of EGUs to identify the viability of biomass co-firing and deep CCS retrofits in the U.S. coal-fired power fleet. Specifically, it characterizes the techno-economic performance of deep carbon capture, estimates life cycle greenhouse gas (GHG) emissions, and conducts a fleet-level assessment on retrofit viability. The key objectives are (1) to estimate the unit-specific performance and retrofitted cost under various biomass co-firing levels and CO2 capture rates; (2) to determine the possibility of reaching net-zero emission at the fleet level; (3) to quantify the cumulative capacities that are suitable for plant retrofits under current and future biomass supply scenarios; and (4) to improve the understanding of policy impacts on such retrofits to help the power sector’s transition to a net-zero economy. Techno-economic Model of Deep Carbon Capture. This study develops the performance and economic models for Monoethanolamine-based post-combustion CO2 capture at 95–99% capture rates. The process is simulated in Aspen Plus, analyzing the performance of carbon capture technology by varying the plant sizes, solvent lean loading, CO2 concentrations, and flue gas inlet temperature. Based on the key inputs and output parameters of CO2 capture, a reduced-order performance model of deep carbon capture is formulated. In addition, an engineering-economic model integrating the performance metrics is developed to estimate the capital as well as operation and maintenance (O&M) costs. Capital cost estimations follow the framework of the Integrated Environmental Control Model (IECM) and incorporate data regressions from three technical reports by IECM, the National Energy Technology Laboratory (NETL), and the National Renewable Energy Laboratory. The O&M cost estimation utilizes the actual inventory consumption rate and labor requirements. Both performance and cost models are embedded into IECM v13.0-beta, a fossil-fuel power plant modeling tool. Life Cycle Assessment of Power Plants. This study estimates the GHG emissions of power plants through life cycle assessment (LCA). The LCA scope includes fuel supply, combustion-based power generation, and CO2 transport and storage. The fuel-based life cycle module is designed following the framework of the NETL Unit Process Library and CO2U LCA Guidance Toolkit. The module is then incorporated into IECM v13.0-beta. The process-based LCA is applied to estimate the GHG emissions of coal and biomass supply, coal- and coal-biomass co-firing power plant operation, as well as CO2 pipeline transport and geographical sequestration. An uncertainty analysis is conducted to quantify the variability and uncertainty associated with the LCA using the Latin Hypercube Sampling (LHS) method. Fleet-level Assessment. This study evaluates the technical and economic feasibility of selected coal-fired EGUs, examines the role of tax credits in retrofit viability, and assesses the competitiveness of retrofitted units against other low-carbon options. Unit screening identifies EGUs for the study, focusing on new, efficient baseload units with air pollution controls. The power plant databases are then established to organize unit-specific information on performance and operating conditions from the relevant public databases. Biomass for co-firing retrofits is selected based on home and neighboring county availability, ensuring sustained operation with at least a 5% co-firing level. The CO2 storage site is determined by state-level storage potential, with ArcGIS Pro and NETL CO2 Saline Storage Cost Model used to identify the optimal balance between the nearest transport distances and affordable storage costs. The latest IECM v13.0-beta is then employed to configure and evaluate the eligible EGUs with or without the deployment of deep CCS and biomass co-firing. A supply curve is established to illustrate the cumulative installed capacity suitable for retrofits at different cost levels. A sensitivity analysis on tax credits for carbon sequestration is performed. Finally, a unit-level cost comparison is conducted among retrofitted plants, renewable power with battery storage, and abated fossil fuels with DAC. Expected Results. This study evaluates the technical, economic, and environmental metrics of each EGU across an array of CO2 capture rates and biomass co-firing level scenarios. Unit-level comparisons will identify critical factors influencing technical performance. The supply curves with and without tax incentives will provide insights into the impact of tax credits on biomass co-firing and CCS deployment. The cost comparisons with renewables and DAC-retrofit will assess the competitiveness of the retrofitted units. Life cycle emissions from each unit will be assessed to identify the scenarios under which net-zero emissions can be achieved. These analyses are expected to determine the total coal-fired capacity suitable for serving as a low-carbon energy source with or without tax incentives. The study results are novel in identifying optimal unit-specific strategies for producing carbon-neutral power, whether through retrofitting EGUs with deep CCS, biomass co-firing, DAC, or installing renewable power with battery. The findings will provide insight into nationwide efforts to ensure reliable, affordable, and low-carbon electricity. It also will inform investment decisions and policies in the deployment of deep carbon capture and negative emission technologies for a net-zero energy future.

Biomass Co-firing↗

HydraGNN_Predictive_GFM_2026 - Ensemble of predictive graph foundation models for atomistic materials modeling

This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0

36 MATERIALS SCIENCE↗

OpenUniverse2024: a shared, simulated view of the sky for the next generation of cosmological surveys

The OpenUniverse2024 simulation suite is a cross-collaboration effort to produce matched simulated imaging for multiple surveys as they would observe a common simulated sky. Both the simulated data and associated tools used to produce it are intended to uniquely enable a wide range of studies to maximize the science potential of the next generation of cosmological surveys. We have produced simulated imaging for approximately 70 deg 2 of the Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST) Wide-Fast-Deep survey and the Nancy Grace Roman Space Telescope High-Latitude Wide-Area Survey, as well as overlapping versions of the ELAIS-S1 Deep-Drilling Field for LSST and the High-Latitude Time-Domain Survey for Roman. OpenUniverse2024 includes (i) an early version of the updated extragalactic model called Diffsky, which substantially improves the realism of optical and infrared photometry of objects, compared to previous versions of these models; (ii) updated transient models that extend through the wavelength range probed by Roman and Rubin; and (iii) improved survey, telescope, and instrument realism based on up-to-date survey plans and known properties of the instruments. It is built on a new and updated suite of simulation tools that improves the ease of consistently simulating multiple observatories viewing the same sky. The approximately 400 TB of synthetic survey imaging and simulated universe catalogs are publicly available, and we preview some scientific uses of the simulations.

large-scale structure of Universe↗

XRISM/Resolve view of Abell 2319: Turbulence, sloshing, and ICM dynamics

Here, we present results from XRISM/Resolve observations of the core of the galaxy cluster Abell 2319, focusing on its kinematic properties. The intracluster medium (ICM) exhibits temperatures of approximately 8 keV across the core, with a prominent cold front and a high-temperature region (⁠~11 keV) in the north-west. The average gas velocity in the 3' x 4' region around the brightest cluster galaxy (BCG) covered by two Resolve pointings is consistent with that of the BCG to within 40 km s -1 ⁠ and we found modest average velocity dispersion of 230–250 km s -1 ⁠. On the other hand, spatially resolved spectroscopy reveals interesting variations. A blueshift of up to ~230 km s -1 ⁠ is observed around the east edge of the cold front, where the gas with the lowest specific entropy is found. The region further south inside the cold front shows only a small velocity difference from the BCG; however, its velocity dispersion is enhanced to ~400 km s -1 ⁠, implying the development of turbulence. These characteristics indicate that we are observing sloshing motion with some inclination angle following BCG and that gas phases with different specific entropy participate in sloshing with their own velocities, as expected from simulations. No significant evidence for a high-redshift ICM component associated with the subcluster Abell 2319B was found in the region covered by the current Resolve pointings. These results highlight the importance of sloshing and turbulence in shaping the internal structure of Abell 2319. Further deep observations are necessary to better understand the mixing and turbulent processes within the cluster.

Astronomy and AstroPhysics↗

Convergent evolution of NFP -facilitated root nodule symbiosis

The origin and phylogenetic distribution of symbiotic associations between nodulating angiosperms and nitrogen-fixing bacteria have long intrigued biologists. Recent comparative evolutionary analyses have yielded alternative hypotheses: a multistep pathway of independent gains and losses of root nodule symbiosis vs. a single gain followed by numerous losses. A detailed reconstruction of the history of genes involved in signaling between nitrogen-fixing bacteria and potential hosts, particularly lipo-chitooligosaccharide (LCO) signaling, is needed to distinguish between these hypotheses. LCO recognition by plants involves the Nod Factor Perception ( NFP ) gene family; in the legume model Medicago truncatula (Fabales), MtNFP is essential for establishing rhizobial symbiosis. Here, we document convergent evolution of NFP , indicating multiple origins of LCO-driven symbiosis. In contrast to previous models that explain the recruitment of NFP via a single duplication in the ancestor of the nitrogen-fixing clade, our phylogenomic and synteny results suggest this duplication does not span the entire clade. Tandem duplication in a common ancestor of Cucurbitales and Rosales resulted in the NFP1 and NFP2 groups. In contrast, the phylogenetically closest paralog of MtNFP is MtLYR1 , located on a different chromosome within a large syntenic block. All available data indicate that a large-scale duplication resulted in MtNFP and MtLYR1 , likely corresponding to a whole-genome duplication in an ancestor of subfamily Papilionoideae of Fabaceae. We show that MtNFP and the NFP2 -like group are not orthologous, indicating multiple independent gains of NFP -based LCO signaling. This molecular convergence provides a possible mechanism for multiple gains of root nodule symbiosis across the nitrogen-fixing clade.

LCO signaling↗

Spectroscopic Characterization of redMaPPer Galaxy Clusters with DESI

Optical galaxy cluster identification algorithms such as redMaPPer promise to enable an array of astrophysical and cosmological studies, but suffer from biases whereby galaxies in front of and behind a galaxy cluster are mistakenly associated with the primary cluster halo. These projection effects caused by irreducible photometric redshift uncertainty must be quantified to facilitate the use of optical cluster catalogues. We present measurements of galaxy cluster projection effects and velocity dispersion using spectroscopy from the Dark Energy Spectroscopic Instrument. Our findings are as follows: we confirm that the fraction of redMaPPer putative member galaxies mistakenly associated with cluster haloes is richness dependent, being more than twice as large at low richness than high richness; we present the first spectroscopic evidence of an increase in projection effects with increasing redshift, by as much as 25 per cent from $z\sim 0.1$ to $z\sim 0.2$; moreover, we find qualitative evidence for luminosity dependence in projection effects, with fainter galaxies being more commonly far behind clusters than their bright counterparts; finally, we fit the scaling relation between measured mean spectroscopic richness and velocity dispersion, finding an implied linear scaling between spectroscopic richness and halo mass. We discuss further directions for the application of spectroscopic data sets to improve use of optically selected clusters to test cosmological models.

clusters↗

Discovery of a nearby radio relic in the low-mass, merging cluster Abell 4067

Shock waves generated during cluster mergers offer a powerful probe of how large-scale structure grows and evolves in the Universe. As part of the MeerKAT-South Pole Telescope (SPT) survey, we report the discovery of a single arc-like radio relic in the galaxy cluster Abell 4067 ($z=0.099$), one of the lowest-mass clusters known to host such a structure. MeerKAT UHF-band (0.58–1.09 GHz) observations reveal a relic with a largest linear size of $\sim 1.48\pm 0.02$ Mpc, located at a projected distance of 0.95 Mpc from the cluster centre. XMM–Newton X-ray data show that the relic’s position and orientation relative to the intracluster medium (ICM) elongation are consistent with a merger-driven shock-wave scenario. The relic has an estimated radio power of $3.10\pm 0.03\times 10^{24}$ W Hz$^{-1}$ at 150 MHz. When placed in the $P_{150\, \mathrm{MHz}}$–$M_{500}$ scaling relation, the Abell 4067 relic appears less luminous compared to relics in more massive clusters, suggesting an association with weak merger shocks. This finding supports the idea that relics in low-mass clusters may form through less energetic merger events, leading to weak merger shocks. The latter is supported by the absence of a detectable central radio halo in Abell 4067, which reinforces the idea that luminous radio haloes are not a universal outcome of cluster mergers and highlights the role of cluster mass, merger energetics, and evolutionary stage in shaping diffuse radio emission in the ICM.

79 ASTRONOMY AND ASTROPHYSICS↗

Proposed New STELLA-2 Tests

In 2025, KAERI and Argonne initiated a collaboration based on software validation using data from tests performed at KAERI’s STELLA-2 facility. STELLA-2 is a large-scale sodium thermal-hydraulic test facility that was originally developed to support KAERI’s development of the PGSFR design concept. KAERI has performed a large number of tests at STELLA-2, providing valuable data to support sodium fast reactor code validation. KAERI will be conducting new tests in 2026, targeting system conditions that were not achieved during previous test campaigns. KAERI has requested from Argonne proposed tests that could be performed during the upcoming test campaign. This report documents Argonne’s proposed tests in fulfillment of KAERI’s request.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Clustering of DESI galaxies split by thermal Sunyaev-Zeldovich effect

The thermal Sunyaev-Zeldovich (tSZ) effect is associated with galaxy clusters - extremely large and dense structures tracing the dark matter with a higher bias than isolated galaxies. We propose to use the tSZ data to separate galaxies from redshift surveys into distinct subpopulations corresponding to different densities and biases independently of the redshift survey systematics. Leveraging the information from different environments, as in density-split and density-marked clustering, is known to tighten the constraints on cosmological parameters, like $\Omega_m$, $\sigma_8$ and neutrino mass. We use data from the Dark Energy Spectroscopic Instrument (DESI) and the Atacama Cosmology Telescope (ACT) in their region of overlap to demonstrate informative tSZ splitting of Luminous Red Galaxies (LRGs). We discover a significant increase in the large-scale clustering of DESI LRGs corresponding to detections starting from 1-2 sigma in the ACT DR6 + Planck tSZ Compton-$y$ map, below the cluster candidate threshold (4 sigma). We also find that such galaxies have higher line-of-sight coordinate (and velocity) dispersions and a higher number of close neighbors than both the full sample and near-zero tSZ regions. We produce simple simulations of tSZ maps that are intrinsically consistent with galaxy catalogs and do not include systematic effects, and find a similar pattern of large-scale clustering enhancement with tSZ effect significance. Moreover, we observe that this relative bias pattern remains largely unchanged with variations in the galaxy-halo connection model in our simulations. This is promising for future cosmological inference from tSZ-split clustering with semi-analytical models. Thus, we demonstrate that valuable cosmological information is present in the lower signal-to-noise regions of the thermal Sunyaev-Zeldovich map, extending far beyond the individual cluster candidates.

Astronomy data analysis↗

Large Eddy Simulation of the Diurnal Cycle of Shallow Convection in the Central Amazon

Climate models often face challenges in accurately simulating the daily precipitation cycle over tropical land areas, particularly in the Amazon. One contributing factor may be the incomplete representation of the diurnal evolution of shallow cumulus (ShCu) clouds. This study aimed to enhance the understanding of the diurnal cycles of ShCu clouds—from formation to maturation and dissipation—over the Central Amazon (CAMZ). Using observational data from the Green Ocean Amazon 2014 (GoAmazon) campaign and large eddy simulation (LES) modeling, we analyzed the diurnal cycles of six selected pure ShCu cases and their composite behavior. Our results revealed a well-defined cycle, with cloud formation occurring between 10 and 11 local time (LT), maturity from 13 to 15 LT, and dissipation by 17–18 LT. The vertical extent of the liquid water mixing ratio and the intensity of the updraft mass flux were closely associated with increases in turbulent kinetic energy (TKE), enhanced buoyancy flux within the cloud layer, and reduced large-scale subsidence. We further analyzed the diurnal cycles of the convective available potential energy (CAPE), the convective inhibition (CIN), the Bowen ratio (BR), and the vertically integrated TKE in the mixed layer (ITKE-ML), exploring their relationships with the cloud base mass flux (Mb) and cloud depth across the six ShCu cases. ITKE-ML and Mb exhibited similar diurnal trends, peaking at approximately 14–15 LT. However, no consistent relationships were found between CAPE (or BR) and Mb. Similarly, comparisons of the cloud depth with CAPE, BR, ITKE-ML, CIN, and Mb revealed no clear relationships. Smaller ShCu clouds were sometimes linked to higher CAPE and lower CIN. It is important to emphasize that these findings are preliminary and based on a limited sample of ShCu cases. Further research involving an expanded dataset and more detailed analyses of the TKE budget and synoptic conditions is necessary. Such efforts would yield a more comprehensive understanding of the factors influencing ShCu clouds’ vertical development.

54 ENVIRONMENTAL SCIENCES↗

ARCTRON: A Rapid Experimental Proving Ground for TPS Experiments and Arcjet Technology Development

Innovation in high-enthalpy facilities is fundamentally limited by the cost and risk of experimentation. New concepts for plasma control, diagnostics, facility components, and plasma-material interaction often require repeated iterations that are impractical to perform in production arcjets. As a result, promising ideas may remain unexplored or reach operational facilities only after significant development effort. ARCTRON is being developed as a rapid experimental proving ground where new ideas in plasma science, arcjet engineering, diagnostics, and material response can be conceived, tested, and quantitatively evaluated before transition to large-scale facilities. The platform combines radio-frequency (RF) and DC arc plasma generation, externally applied magnetic fields, configurable gas composition, reduced-pressure operation, laser heating, electrical biasing, and modular diagnostic access. These capabilities permit the plasma source, applied forcing, test article, and measurement configuration to be modified independently, allowing individual physical mechanisms to be isolated more readily than in a traditional test environment. One class of investigations addresses fundamental plasma-surface interaction physics. Conventional material tests often expose a specimen simultaneously to convective heating, reactive species, pressure, shear, radiation, and surface-current effects. The resulting material response may be measured accurately, while the contribution of each mechanism remains difficult to identify. ARCTRON is designed to vary these effects selectively. Plasma chemistry can be changed independently through configurable gas mixtures; magnetic fields and electrical biasing can modify charged-particle transport; laser heating can provide a non-plasma thermal input; and pressure, flow, and discharge mode can be varied over a broad operating space. This enables controlled tests of hypotheses involving surface catalycity, reactive-species transport, plasma-assisted oxidation, electromagnetic effects, shear, and the relative contributions of thermal and chemical loading. A second class of investigations enabled by this approach concerns the engineering of high-enthalpy facilities themselves. Arc-heated facilities are limited by electrode erosion, unstable arc attachment, localized heating, and damage to nozzles and other plasma-facing components. ARCTRON provides a lower-cost environment for testing concepts intended to mitigate these limitations. Candidate investigations include the use of applied magnetic fields to alter current paths and reduce plasma interaction with nozzle walls, ExB forcing to introduce controlled plasma rotation, magnetic or geometric approaches for distributing arc attachment, and alternative electrode or discharge configurations intended to reduce erosion and improve stability. Because the platform is reconfigurable, these concepts can be evaluated through repeated design--build--test cycles before they are considered for implementation in operational facilities. The platform also supports the development and validation of diagnostics that may be difficult to introduce initially into a large arcjet. Current and planned measurements include spatially resolved optical emission spectroscopy, electrostatic probes, fast imaging, pyrometry, calorimetry, laser-induced fluorescence, and absorption spectroscopy. These diagnostics are intended not merely to document a nominal operating condition, but to constrain the local plasma state and its relationship to component or material response. The modular facility geometry allows diagnostic concepts to be tested, calibrated, and compared under repeatable conditions before deployment in more demanding environments. ARCTRON is also supported by an integrated software suite. Automated control and data acquisition allow discharge parameters, gas composition, magnetic fields, diagnostic timing, and test configuration to be recorded as part of each experiment (STARDAC - Software for Testing, Analysis, Research Data, and Control). The Backend for Experiment Analysis, Storage, and Traceability (BEAST) is a database that provides the infrastructure needed to associate heterogeneous measurements with facility configuration, specimen identity, calibration state, geometry, and analysis provenance. This backend is particularly important for exploratory campaigns, in which many related configurations may be tested, and the value of an individual experiment depends on its connection to earlier and subsequent iterations. Complementary analysis capabilities, including computer-vision-based transient response measurements (arcjetCV), three-dimensional surface reconstruction (STARSCAN), and model-based Bayesian inference (SHIELD), and tomography data analysis (TOMATO, PuMA) can be incorporated when required by a specific hypothesis without becoming the focus of every campaign. The central objective of ARCTRON is therefore not to maximize heat flux or reproduce a complete flight environment. Its purpose is to reduce the cost and time required to ask consequential questions about plasma behavior, plasma-facing materials, diagnostics, and arcjet technology. By providing a controlled environment for rapid reconfiguration, mechanism isolation, quantitative measurement, and iterative engineering, ARCTRON can help mature concepts that would otherwise remain too speculative or too risky for evaluation in production facilities. The resulting knowledge can then guide the design of material models, focus test objectives in larger arcjets, reduce facility-development risk, and improve the physical basis of high-enthalpy ground testing. This work will present the ARCTRON architecture, operating modes, diagnostic suite, and digital experimental workflow. Initial experimental results from the first integrated operation of the facility will be presented, including flow characterization, power limitations, and deployment of the initial diagnostic suite. Ongoing development efforts aimed at catalycity characterization, magnetic plasma control, and advanced optical diagnostics will also be discussed, illustrating how the platform supports rapid iteration from concept to experiment.

experimental diagnostics↗

ARCTRON: A Rapid Experimental Proving Ground for TPS Experiments and Arcjet Technology Development

Innovation in high-enthalpy facilities is fundamentally limited by the cost and risk of experimentation. New concepts for plasma control, diagnostics, facility components, and plasma-material interaction often require repeated iterations that are impractical to perform in production arcjets. As a result, promising ideas may remain unexplored or reach operational facilities only after significant development effort. ARCTRON is being developed as a rapid experimental proving ground where new ideas in plasma science, arcjet engineering, diagnostics, and material response can be conceived, tested, and quantitatively evaluated before transition to large-scale facilities. The platform combines radio-frequency (RF) and DC arc plasma generation, externally applied magnetic fields, configurable gas composition, reduced-pressure operation, laser heating, electrical biasing, and modular diagnostic access. These capabilities permit the plasma source, applied forcing, test article, and measurement configuration to be modified independently, allowing individual physical mechanisms to be isolated more readily than in a traditional test environment. One class of investigations addresses fundamental plasma-surface interaction physics. Conventional material tests often expose a specimen simultaneously to convective heating, reactive species, pressure, shear, radiation, and surface-current effects. The resulting material response may be measured accurately, while the contribution of each mechanism remains difficult to identify. ARCTRON is designed to vary these effects selectively. Plasma chemistry can be changed independently through configurable gas mixtures; magnetic fields and electrical biasing can modify charged-particle transport; laser heating can provide a non-plasma thermal input; and pressure, flow, and discharge mode can be varied over a broad operating space. This enables controlled tests of hypotheses involving surface catalycity, reactive-species transport, plasma-assisted oxidation, electromagnetic effects, shear, and the relative contributions of thermal and chemical loading. A second class of investigations enabled by this approach concerns the engineering of high-enthalpy facilities themselves. Arc-heated facilities are limited by electrode erosion, unstable arc attachment, localized heating, and damage to nozzles and other plasma-facing components. ARCTRON provides a lower-cost environment for testing concepts intended to mitigate these limitations. Candidate investigations include the use of applied magnetic fields to alter current paths and reduce plasma interaction with nozzle walls, ExB forcing to introduce controlled plasma rotation, magnetic or geometric approaches for distributing arc attachment, and alternative electrode or discharge configurations intended to reduce erosion and improve stability. Because the platform is reconfigurable, these concepts can be evaluated through repeated design--build--test cycles before they are considered for implementation in operational facilities. The platform also supports the development and validation of diagnostics that may be difficult to introduce initially into a large arcjet. Current and planned measurements include spatially resolved optical emission spectroscopy, electrostatic probes, fast imaging, pyrometry, calorimetry, laser-induced fluorescence, and absorption spectroscopy. These diagnostics are intended not merely to document a nominal operating condition, but to constrain the local plasma state and its relationship to component or material response. The modular facility geometry allows diagnostic concepts to be tested, calibrated, and compared under repeatable conditions before deployment in more demanding environments. ARCTRON is also supported by an integrated software suite. Automated control and data acquisition allow discharge parameters, gas composition, magnetic fields, diagnostic timing, and test configuration to be recorded as part of each experiment (STARDAC - Software for Testing, Analysis, Research Data, and Control). The Backend for Experiment Analysis, Storage, and Traceability (BEAST) is a database that provides the infrastructure needed to associate heterogeneous measurements with facility configuration, specimen identity, calibration state, geometry, and analysis provenance. This backend is particularly important for exploratory campaigns, in which many related configurations may be tested, and the value of an individual experiment depends on its connection to earlier and subsequent iterations. Complementary analysis capabilities, including computer-vision-based transient response measurements (arcjetCV), three-dimensional surface reconstruction (STARSCAN), and model-based Bayesian inference (SHIELD), and tomography data analysis (TOMATO, PuMA) can be incorporated when required by a specific hypothesis without becoming the focus of every campaign. The central objective of ARCTRON is therefore not to maximize heat flux or reproduce a complete flight environment. Its purpose is to reduce the cost and time required to ask consequential questions about plasma behavior, plasma-facing materials, diagnostics, and arcjet technology. By providing a controlled environment for rapid reconfiguration, mechanism isolation, quantitative measurement, and iterative engineering, ARCTRON can help mature concepts that would otherwise remain too speculative or too risky for evaluation in production facilities. The resulting knowledge can then guide the design of material models, focus test objectives in larger arcjets, reduce facility-development risk, and improve the physical basis of high-enthalpy ground testing. This work will present the ARCTRON architecture, operating modes, diagnostic suite, and digital experimental workflow. Initial experimental results from the first integrated operation of the facility will be presented, including flow characterization, power limitations, and deployment of the initial diagnostic suite. Ongoing development efforts aimed at catalycity characterization, magnetic plasma control, and advanced optical diagnostics will also be discussed, illustrating how the platform supports rapid iteration from concept to experiment.

experimental diagnostics↗