Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Large-Scale Scientific Simulation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Feeding Ten Billion People Is Possible Within Four Terrestrial Planetary Boundaries

Global agriculture puts heavy pressure on planetary boundaries, posing the challenge to achieve future food security without compromising Earth system resilience. On the basis of process-detailed, spatially explicit representation of four interlinked planetary boundaries (biosphere integrity, land-system change, freshwater use, nitrogen flows) and agricultural systems in an internally consistent model framework, we here show that almost half of current global food production depends on planetary boundary transgressions. Hotspot regions, mainly in Asia, even face simultaneous transgression of multiple underlying local boundaries. If these boundaries were strictly respected, the present food system could provide a balanced diet (2,355 kcal per capita per day) for 3.4 billion people only. However, as we also demonstrate, transformation towards more sustainable production and consumption patterns could support 10.2 billion people within the planetary boundaries analysed. Key prerequisites are spatially redistributed cropland, improved water–nutrient management, food waste reduction and dietary changes. Adoption of the Sustainable Development Goals by all nations in 2015 is the first ever commitment to a world development path that safeguards the stability of the Earth system as a prerequisite for meeting universal human standards1. The longstanding challenge of achieving food security through sustainable agriculture is particularly acute in this context as world agriculture is a leading cause for the current transgressions of multiple planetary boundaries (PBs) globally and regionally2–5. The PB framework is a comprehensive scientific attempt to synoptically define our planet’s biogeophysical limits to anthropogenic interference. It suggests bounds to nine interacting processes that together delineate a Holocene-like Earth system state. The Holocene is chosen as the reference state as it is the only period known to provide a safe operating space for a world population of several billion people, and according to a precautionary principle, the PBs are set in sufficient distance from processes that may critically undermine Earth system resilience and global sustainability. A challenging question, thus, is whether human development goals such as food security can be met while maintaining multiple PBs along with their subglobal manifestations. Further PB transgressions could jeopardize the chances of providing sufficient food for a world population projected to be wealthier and reach >9 billion by 2050. This conundrum portrays a tradeoff between Earth’s biophysical carrying capacity and humankind’s rising food demand, calling in response for radical rethinking of food production and consumption patterns6–9. Yield gap closures, avoidance of excessive input use, shifts towards less resource-demanding diets, food waste reductions and efficient international trade are crucial options for sustainably increasing the food supply10–15. For example, enhancing water-use efficiency on irrigated and rain-fed farms can triple or quadruple crop yields in low-performing systems, suggesting possible global gains of >20% (ref. 16). Even higher gains appear feasible through globally optimized configurations of the land-use pattern17, and cutting food losses by half could generate food for another billion people18. Thus, collective large-scale implementation of such options could sustain food for a further growing world population19. Yet achieving this within a safe operating space as defined by PBs requires not only a halt to but actually a reversal of existing PB transgressions. Previous studies suggest that such a reconciliation might be possible, but these were based on aggregate representations of PBs (not accounting for the spatial patterns of limits, transgressions and interactions) or considered only one boundary in isolation17,20–23. Here, we systematically quantify to what extent current food production depends on local to global transgressions of the PBs for biosphere integrity, land-system change, freshwater use and nitrogen (N) flows, along with the potential of a range of solutions to avoid these transgressions and still increase food supply (Table 1). To this end, we configured an internally consistent process-based model of the terrestrial biosphere including agriculture (LPJmL) with multiple spatially distributed PBs and their interactions. LPJmL is among the longest-established and best-evaluated biosphere models, showing robust performance regarding simulation of, for example, carbon, water and crop yield dynamics (Supplementary Figs. 1 and 2 and Supplementary Table 1; see ref. 24 for a comprehensive benchmarking and Supplementary Methods for more detail on model evaluations). In principle following established definitions4, we refine the computation of some PBs with respect to their regional patterns and interactions (Methods), providing globally gridded precautionary limits to human interference with the Earth system at a level of great detail. In particular, we account for the evidence that many PBs need to be represented spatially explicitly4 to cover their

Gerten, Dieter↗

2D reactive transport model of shale chemical weathering and biogeochemical fluxes along a mountainous hillslope, East River Watershed, Colorado: Input files and simulation results

This data package contains input files and simulation results for a two-dimensional (2D) reactive transport model used to quantitatively analyze the coupled hydrological and biogeochemical processes governing shale weathering and associated biogeochemical fluxes under realistic environmental conditions in the high-elevation East River Watershed. These data support the conclusions presented in Stolze et al. (Water Resources Research, under review), "Model-based interpretation of solute exports and carbon partitioning during shale weathering in a mountainous hillslope". The model simulates atmospheric-subsurface gas exchange, subsurface water flow, and shale weathering processes under dynamic, year-scale conditions along a shale-underlain hillslope located in the East River watershed. The simulations were performed using the PFLOTRAN flow and reactive transport code and executed on the Perlmutter supercomputer to leverage its large-scale parallel computing capabilities. The data package contains two zipped folders, "model_input_files" and "simulation_results", and one readme.txt file. "model_input_files" contains the necessary input files to run the calibrated base-base model presented in Stolze et al. (Water Resources Research, under review). "simulation_results" contains a single hdf5 file ("Output_2D_hillslope_model.h5") which includes the results of simulation performed using the base-case model. This file can be opened with HDFView 3.1.4, Python, or MATLAB. "readme.txt" contains relevant information about the base-case model and provides guidelines on how to run the associated input files provided in the folder "model_input_files". Furthermore, readme.txt provides information regarding the model results provided in "Output_2D_hillslope_model.h5" such as matrix dimensionality and output units. Field datasets used to evaluate model performance were collected at three monitoring wells located along a hillslope transect (PLM1, PLM2, and PLM3). Dissolved ion concentration data were collected from November 2016 to October 2021 for Ca, Mg, DIC, Na, K, SO4 (Dong et al., 2025 - dic_npoc_data_2014_2024.zip - DOI:10.15485/1660459; Williams et al., 2025 - anion_data_2014_2024.zip - DOI:10.15485/1668054; Dong et al., 2025 - cation_data_2014_2024.zip - DOI:10.15485/1668055). Note that we used the files named er_PLM1_xx_yy, er_PLM2_xx_yy, and er_PLM3_xx_yy where xx stands for the name of the aqueous species and yy stands for the depth where the measurements were performed. Soil water content ([0 - 1] m) and water table depth were collected from November 2016 to October 2021 (Wan et al., 2024 - Dynamic_water_table__depthsFig2b.csv and Soil_water_content_Fig4e.csv - DOI:10.15485/2322567). Gaseous CO2 concentration were collected from October 2020 to December 2021(Wan et al., 2024 - Soil_CO2_concentrations_Fig4h.csv - DOI:10.15485/2322567) Gaseous CO2 flux from the subsurface to the atmosphere were collected in the vicinity of PLM2 from October 2019 to May 2022 (Wu et al., 2025). Soil microbial biomass concentration was measured from August 2016 to June 2017 (Sorensen et al., 2019 - 2017_East_River_Pumphouse_Microbial_Biomass__1_.csv - DOI:10.15485/1577267) All field data are published as CSV files compatible with Microsoft Excel, MATLAB, and Python, or as text files. The coordinates of the monitoring wells and the CO2(g) flux sensor in the coordinate system WGS84 are: -PLM1: [38.9197710 ; -106.9492750] -PLM2: [38.9201580 ; -106.9487170] -PLM3: [38.9207843 ; -106.9483668] -PLM4: 38.9210060 ; -106.9479528] -CO2(g) flux sensor: [38.9199180 ; -106.9489906] ------------------------------------------------------------------------------------------- This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. This research used resources of the National Energy Research Scientific Computing Center (NERSC), a Department of Energy User Facility using NERSC award BER-ERCAP 23980, BER-ERCAP 28550, and BER-ERCAP 33789.

54 ENVIRONMENTAL SCIENCES↗

Cosmology with the Roman Space Telescope - Multiprobe Strategies

We simulate the scientific performance of the Nancy Grace Roman Space Telescope High Latitude Survey (HLS) on dark energy and modified gravity. The 1.6-yr HLS Reference survey is currently envisioned to image 2000 deg2 in multiple bands to a depthof∼26.5 in Y, J, H and to cover the same area with slit-less spectroscopy beyond z=3. The combination of deep, multiband photometry and deep spectroscopy will allow scientists to measure the growth and geometry of the Universe through a variety of cosmological probes (e.g. weak lensing, galaxy clusters, galaxy clustering, BAO, Type Ia supernova) and, equally, it will allow an exquisite control of observational and astrophysical systematic effects. In this paper, we explore multiprobe strategies that can be implemented, given the telescope’s instrument capabilities. We model cosmological probes individually and jointly and account for correlated systematics and statistical uncertainties due to the higher order moments of the density field. We explore different levels of observational systematics for the HLS survey (photo-z and shear calibration) and ultimately run a joint likelihood analysis in N-dim parameter space. We find that the HLS reference survey alone can achieve a standard dark energy FoM of>300 when including all probes. This assumes no information from external data sets, we assume a flat universe however, and includes realistic assumptions for systematics. Our study of the HLS reference survey should be seen as part of a future community-driven effort to simulate and optimize the science return of the Roman Space Telescope.

Cosmological parameters↗

OpenUniverse2024: a shared, simulated view of the sky for the next generation of cosmological surveys

The OpenUniverse2024 simulation suite is a cross-collaboration effort to produce matched simulated imaging for multiple surveys as they would observe a common simulated sky. Both the simulated data and associated tools used to produce it are intended to uniquely enable a wide range of studies to maximize the science potential of the next generation of cosmological surveys. We have produced simulated imaging for approximately 70 deg 2 of the Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST) Wide-Fast-Deep survey and the Nancy Grace Roman Space Telescope High-Latitude Wide-Area Survey, as well as overlapping versions of the ELAIS-S1 Deep-Drilling Field for LSST and the High-Latitude Time-Domain Survey for Roman. OpenUniverse2024 includes (i) an early version of the updated extragalactic model called Diffsky, which substantially improves the realism of optical and infrared photometry of objects, compared to previous versions of these models; (ii) updated transient models that extend through the wavelength range probed by Roman and Rubin; and (iii) improved survey, telescope, and instrument realism based on up-to-date survey plans and known properties of the instruments. It is built on a new and updated suite of simulation tools that improves the ease of consistently simulating multiple observatories viewing the same sky. The approximately 400 TB of synthetic survey imaging and simulated universe catalogs are publicly available, and we preview some scientific uses of the simulations.

large-scale structure of Universe↗

Investigating Low-Altitude Constellations of Ad-Hoc Lunar PNT System for Distributed Spacecraft Autonomy

In this study, we examine a low-altitude Lunar Position, Navigation, and Timing (LPNT) constellations and the localization performance of Centralized Extended Kalman Filter (CEKF) and Decentralized Extended Kalman Filter (DEKF) algorithms. The primary investigation involves a 100-node swarm operating at a 100 km altitude, in contrast to previous studies that examined a 21-node asset in a frozen-orbit at 5,500 km. The autonomous operation of large-scale swarm is based on two-way Inter-Satellite Link (ISL) measurements, which involve pseudoranges and relative velocities among swarm nodes. We perform a numerical assessment of the two filtering approaches, utilizing ‘fully sampled’ measurements from all available assets as well as ‘two ISL’ measurements where each spacecraft is restricted to only two antennas. This research includes an analysis of CEKF under 2-ISL constraints and evaluates the performance of DEKF in a 100-node swarm, which has not been explored in previous studies. In addition, we examine the impact of increasing the sampling frequency for DEKF, showing that the update cycle can be shortened from a 10-minute interval. A novel approach for ‘2-ISL limited’ DEKF will also be introduced, using a matching formulation that exhaustively enumerates all potential matches. This study provides valuable insights into large-scale distributed swarm operations, considering various filter configurations, sampling frequencies, matching strategies, and scalability of CEKF and DEKF for low-altitude LPNT applications. The Lunar PNT technology plays a key role in providing reliable and robust navigation services on the Moon's surface and the South pole, where the primary Lunar missions are planned. To support upcoming Lunar missions, including small satellites from NASA's Commercial Lunar Payload Services program, the Lunar PNT system must be adaptable to smaller platforms like CubeSats. Driven by the growing involvement of public and private exploration partnerships, the traditional low Earth orbit missions are shifting to beyond geosynchronous orbit [1]. These upcoming missions aim to foster a sustainable and innovative exploration program, in collaboration with commercial and international partners, to facilitate human expansion throughout the solar system and return new knowledge and opportunities to Earth [2]. As part of this trend, there are increasing efforts to utilize science missions in Lunar orbit to develop a non-dedicated and ad-hoc PNT network system. Two traditional approaches, the Deep Space Network (DSN) and the weak signal Global Positioning System (GPS), are established deep-space navigation technologies for missions beyond the geosynchronous orbit. Beginning in 1958, the DSN was developed to communicate with the Explorer 1 spacecraft based on the use of radiometric tracking in spacecraft navigation [3]. The DSN is capable of providing nearly unfettered coverage to spacecraft beyond low-Earth orbit (LEO), however, increased space mission volume has created concerns about future expectations of DSN usage for spacecraft navigation [4]. For cislunar mission applications, the position accuracy using DSN achieves 100 m (3σ) with at least three geometrically diverse ground stations when using radiometric tracking alone [5]. The DSN's dependence on Earth-based ground stations restricts its operational capabilities to periods of Earth visibility. This limitation, coupled with its poor localization performance, renders the DSN unsuitable for future lunar missions that demand continuous tracking and precise positioning. To satisfy the increasing requirements of DSN in Lunar applications, spacecrafts are also required to improve their onboard antenna power and efficiency of the transmission. However, there is an important aggregate cost trade between adding capabilities to every spacecraft and adding to a capacity on the ground that serves multiple spacecraft [6]. A weak GPS system can provide PNT service while the user spacecraft is bound to the Moon, leveraging a single, steerable high gain antenna with the relatively narrow beam which includes all the sources in its field of view [7]. However, the higher the altitude the receiver is above the GPS constellations, the poorer and the weaker are the relative geometry and the received signal powers, respectively, leading to a significant navigation accuracy reduction [8]. The transmitted power becomes weaker with increasing distance from the Earth as well as signals tracked from one of the side lobes of the GPS antenna pattern. As a results, the number of visible satellites and relative geometric condition of the GPS satellites at very high altitude drops dramatically and reduces the navigation solution accuracy. Therefore, the weak GPS system is also not an ideal way to provide PNT service to upcoming Lunar missions when considering its limited geometric condition and the recued navigation accuracy. Another navigation approach on the Moon is being developed, similar to the Global Navigation Satellite System (GNSS) on Earth, aiming to offer navigation service with continuous 24/7 coverage across the entire Lunar surface. For example, lunar communications relay and navigation systems (LCRNS) by NASA and Lunar navigation satellite systems (LNSS) by JAXA are designed to serve as dedicated Position, Navigation, and Timing (PNT) systems for the Moon. However, designing a dedicated LNSS and PNT service involves additional challenges, which are unique to the lunar environment, including limited payload capacity for the CubeSat platform, i.e., the size, weight, and power (SWaP) of the onboard clock, limited lunar ground monitoring stations, and limited financial investment as compared to the legacy Earth-GPS [9]. NASA’s focus on utilizing CubeSat platforms on the Moon leads to an alternative Lunar navigation platform that leverages the existing Lunar science and exploration assets. The small satellites used in Lunar missions can be used to create a low-cost, autonomous, ad-hoc, and on-demand mission-centric Lunar PNT swarm capable of providing PNT services to these low-cost lunar missions [10]. As upcoming Lunar missions will often operate at low-altitude about 30 km to 100 km for scientific observations and mapping purposes, the low-altitude orbital constellations could be employed to create an ad-hoc Lunar PNT system. However, several issues must be addressed, such as the instability of these orbits, which often require maintenance or are only suitable for short-duration missions, operating for fewer than 90 days. Additionally, at an altitude of 100 km, the satellites have a limited period during which they are above the horizon and capable of providing PNT service to users. The implementation of a non-dedicated, ad-hoc Lunar navigation constellation facilitates on-demand PNT services. A preliminary study of ad-hoc Lunar PNT system was conducted using 21 spacecraft in 5,5000 km altitude frozen orbits to test its feasibility and a basic performance of orbital asset localization among ad-hoc Lunar constellations in small satellites format [10]. These swarm assets are designed for autonomous localization with minimal Earth interaction, reducing dependency on bandwidth and ground resources. The design in [10] demonstrated the feasibility of a decentralized PNT approach, specifically employing a DEKF approach for state estimation, which helps minimize onboard operating costs. The DEKF method distributes computation across individual satellites, which lightens the computational load while maintaining accuracy in orbit ephemeris and clock offsets, similar to centralized systems [11]. In a follow-on study [12], each spacecraft was limited to 2 communications antennae, forcing the selection of measurements and scheduling spacecraft activities to perform the measurements. A matching algorithm is implemented to select the best measurements and schedule position estimation updates. The decentralized localization performance is also investigated with increasing levels of network degradation for swarm assets considering the impact of intermittent and permanent communication failure, to demonstrate the robustness and fidelity of the decentralized Lunar PNT service [13]. This study confirmed that the ad-hoc PNT constellations in frozen orbit are highly robust and resilient to communication failures. However, unlike frozen orbit swarm assets, the low-altitude satellites have a limited ground view at an altitude of 100 km, where the ad-hoc Lunar constellation consists of 98 low-altitude satellites, evenly distributed across seven circular polar orbital planes, alongside two satellites in a frozen orbit at an altitude of 5,500 km (Figure 1). Therefore, the number of satellites visible to ground users is significantly limited in low-altitude orbit constellations. As each visibility of a spacecraft remains intact for only a few ticks before it moves out of the field of view, the ground user encounters challenges in maintaining continuous navigation service, resulting in sparse availability and provision of Lunar PNT system. Consequently, service availability is primarily restricted to the Lunar South Pole region (Figure 2). Given these limitations and concerns, the localization performance of low-altitude swarm assets will be assessed in this study. We focus on the investigation of the localization performance of low-altitude swarm assets and ground users near the Lunar South Pole. The overall flow of the Lunar PNT simulation incorporates the DEKF approach of asset localization and the weighted least-squares approach in user localization (Figure 3). The autonomous Lunar PNT simulation is primarily implemented in MATLAB, where the DEKF based on the matching scheduler is implemented with Google’s OR-tools as a model builder and Gurobi optimization tool as a backend solver. The General Mission Analysis Tool (GMAT) is utilized to generate ephemeris data for swarm assets, and accounts for satellite orbital details, mass, and perturbations like solar radiation pressure and drag coefficients. Each ephemeris dataset is produced in the Moon International Celestial Reference Frame (ICRF) inertial coordinate system. For state estimation, the distributed swarm assets rely on two-way Inter-Satellite Link (ISL) measurements, which involve tracking pseudoranges and relative velocities between visible satellites and anchor nodes during each observation. Numerical evaluations of the decentralized localization process are conducted to demonstrate the feasibility of the low-altitude PNT system in providing reliable navigation services. The main approach involves using DEKF and CEKF to localize 100 satellites in low-altitude constellations, where the CEKF is implemented to serve as a baseline for comparing the performance of distributed algorithms. In both cases, we evaluate ‘fully sampled’ measurements from all available assets, and ‘two ISL’ measurements when spacecraft are constrained to have only two antennas. We test four estimation techniques: CEKF fully sampled, CEKF two ISL, DEKF fully sampled, and DEKF two ISL filters. As the DEKF update cycle is comprised of network setup, communication, and computations, a global broadcast network and 2-way ISL network setup will take from 4 to 6 minutes as maximum [12]. In this simulation, the DEKF update cycle is set to 10 minutes, including a 4-minute latency for obtaining and computing the actual measurement updates. We experiment an increased update cycle to demonstrate the feasibility and evaluate the impact on localization performance using various tuning values for measurement noise covariances (Figures 4 and 5). By comparing centralized and decentralized approaches using a matching algorithm, we analyze the influence of cross-correlation factors in the covariance matrix, assuming 100% reliability of all assets and measurements. The increased frequency and the adjustments of tuning parameters reveal distinct error patterns between the two scenarios. The localization accuracy of the swarm assets and ground users is assessed by taking the median error across 100 assets and one ground user (84.9°S, 137.5°E) over 7-day simulation period (Table 1). Since the user localization accuracy is significantly affected by the performance of the swarm assets, it is crucial to maintain high localization accuracy within the swarm. This study will continue to explore decentralized filtering for autonomous LPNT operations, with further investigation of an 'iterative' matching approach which enumerates every valid matching pair, planned for the following month.

Yeji Kim↗

Train small, model big: Scalable physics simulators via reduced order modeling and domain decomposition

Numerous cutting-edge scientific technologies originate at the laboratory scale, but transitioning them to practical industry applications is a formidable challenge. Traditional pilot projects at intermediate scales are costly and time-consuming. An alternative, the pilot-scale model, relies on high-fidelity numerical simulations, but even these simulations can be computationally prohibitive at larger scales. To overcome these limitations, we propose a scalable, physics-constrained reduced order model (ROM) method. The ROM identifies critical physics modes from small-scale unit components, projecting governing equations onto these modes to create a reduced model that retains essential physics details. We also employ Discontinuous Galerkin Domain Decomposition (DG-DD) to apply ROM to unit components and interfaces, enabling the construction of large-scale global systems without data at such large scales. Here this method is demonstrated on the Poisson and Stokes flow equations, showing that it can solve equations about 15–40 times faster with only ~1% relative error. Furthermore, ROM takes one order of magnitude less memory than the full order model, enabling larger scale predictions at a given memory limitation.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

CG-Kit: Code Generation Toolkit for performant and maintainable variants of source code applied to Flash-X hydrodynamics simulations

CG-Kit is a new Code Generation tool-Kit that we have developed as a part of the solution for portability and maintainability for multiphysics computing applications. The development of CG-Kit is rooted in the urgent need created by the shifting landscape of high-performance computing platforms and the algorithmic complexities of a particular large-scale multiphysics application: Flash-X. To efficiently use computing resources on a heterogeneous node, an application must have a map of computation to resources and a mechanism to move the data and computation to the resources according to the map. Most existing performance portability solutions are focussed on abstracting the expression of computations so that a unified source code can be specialized to run on different resources. However, such an approach is insufficient for a code like Flash-X, which has a multitude of code components that can be assembled in various permutations and combinations to form different instances of applications. Similar challenges apply to any code that has composability, where a single specified way of apportioning work among devices may not be optimal. Additionally, use cases arise where the optimal control flow of computation may differ for different devices while the underlying numerics remain identical. This combination leads to unique challenges including handling an existing large code base in Fortran and/or C/C++, subdivision of code into a great variety of units supporting a wide range of physics and numerical methods, different parallelization techniques for distributed and shared memory systems and accelerator devices, and heterogeneity of computing platforms requiring coexisting variants of parallel algorithms. All of these challenges demand that scientific software developers apply existing knowledge about domain applications, algorithms, and computing platforms to determine custom abstractions and granularity for code generation. There is a critical lack of tools to tackle those problems. CG-Kit is designed to fill this gap by providing a user with the ability to express their desired control flow and computation-to-resource map in the form a pseudocode-like recipe. It consists of standalone tools that can be combined into highly specific and, we argue, highly effective portability and maintainability toolchains. Here we present the design of our new tools: parametrized source trees, control flow graphs, and recipes. The tools are implemented in Python. They are agnostic to the programming language of the source code targeted for code generation. In conclusion, we demonstrate the capabilities of the toolkit with two examples, first, multithreaded variants of the basic AXPY operation, and second, variants of parallel algorithms within a hydrodynamics solver, called Spark, from Flash-X that operates on block-structured adaptive meshes.

Algorithmic portability↗

Optical neural engine for solving scientific partial differential equations

Abstract Solving partial differential equations (PDEs) is the cornerstone of scientific research and development. Data-driven machine learning (ML) approaches are emerging to accelerate time-consuming and computation-intensive numerical simulations of PDEs. Although optical systems offer high-throughput and energy-efficient ML hardware, their demonstration for solving PDEs is limited. Here, we present an optical neural engine (ONE) architecture combining diffractive optical neural networks for Fourier space processing and optical crossbar structures for real space processing to solve time-dependent and time-independent PDEs in diverse disciplines, including Darcy flow equation, the magnetostatic Poisson’s equation in demagnetization, the Navier-Stokes equation in incompressible fluid, Maxwell’s equations in nanophotonic metasurfaces, and coupled PDEs in a multiphysics system. We numerically and experimentally demonstrate the capability of the ONE architecture, which not only leverages the advantages of high-performance dual-space processing for outperforming traditional PDE solvers and being comparable with state-of-the-art ML models but also can be implemented using optical computing hardware with unique features of low-energy and highly parallel constant-time processing irrespective of model scales and real-time reconfigurability for tackling multiple tasks with the same architecture. The demonstrated architecture offers a versatile and powerful platform for large-scale scientific and engineering computations.

Tang, Yingheng (ORCID:0009000153622546)↗

Mapping the Perseus galaxy cluster with XRISM: Gas kinematic features and their implications for turbulence

We present extended gas kinematic maps of the Perseus cluster based on a combination of five new XRISM/Resolve pointings observed in 2025 with four performance verification datasets from 2024, totaling a net exposure of 745 ks. To date, Perseus remains the only cluster that has been extensively mapped out to ≃0.7 r 2500 by XRISM/Resolve, while simultaneously offering sufficient spatial resolution to resolve gaseous substructures driven by mergers and active galactic nucleus (AGN) feedback. Our observations cover multiple radial directions and a broad range of dynamical scales, enabling us to characterize the kinematic properties of the intracluster medium up to a scale of ∼500 kpc. In the measurements, we detected high-velocity dispersions (≃300km s −1 ) in the eastern region of the cluster that are spatially coincident with the extended X-ray surface brightness excess and correspond to a nonthermal pressure fraction of ≃7 − 13%. The velocity field outside the AGN-dominant region can be effectively described by a single, large-scale kinematic driver based on the velocity structure function, which statistically favors an energy injection scale of at least a few hundred kpc. The estimated turbulent dissipation energy is comparable to the gravitational potential energy released by a recent merger, implying a significant role of turbulent cascade in the merger energy conversion. In the bulk velocity field, we observed a dipole-like pattern along the east-west direction with an amplitude of ≃ ± 200 − 300 km s −1 , indicating rotational motions induced by the recent merger event. This feature constrains the viewing direction to ≃30° −50° relative to the normal of the merger plane. Our hydrodynamic simulations suggest that Perseus has experienced at least two energetic mergers since redshift z ∼ 1, the most recent of which is associated with the radio galaxy IC310, in agreement with recent SRG/eROSITA findings. This study showcases exciting scientific opportunities for future missions with high-resolution spectroscopic capabilities (e.g., HUBS, LEM, and NewAthena).

79 ASTRONOMY AND ASTROPHYSICS↗

The Origins of the ParaView and VisIt Scientific Visualization Tools

ParaView and VisIt play a key role in the visual understanding of scientific simulation data. These tools are open source, designed to handle extremely large datasets, and can run on supercomputers and large-scale display walls. They are in daily use by scientists, practitioners, and students at supercomputing centers, in industry, and at universities. In conclusion, this article gives personal accounts of the origins of these visualization tools.

Ahrens, James [Los Alamos National Laboratory (LAN↗

To Derive or Not to Derive: I/O Libraries Take Charge of Derived Quantities Computation

The ever-increasing volume of data produced by HPC simulations necessitates scalable methods for data exploration and knowledge extraction. Scientific data analysis often involves complex queries across distributed datasets, requiring manipulation of multiple primary variables and generating derived data that needs to be handled efficiently, creating challenges for applications that need to parse many large datasets. Relying on individual applications to handle all intermediate data generally leads to redundant computations across studies and unnecessary data transfers. In this paper, we investigate the performance of different approaches where applications define derived variables as quantities of interest (QoIs) and offload the computation and transfer of these QoIs to the I/O library. This significantly reduces redundancy and optimizes data movement across the distributed storage and processing infrastructure by allowing control over when and where derived variables are computed. We present a detailed analysis of the performance-storage trade-offs associated with different solutions and showcase results for our study on two large-scale datasets created from climate and combustion simulations.

Gainaru, Ana↗

High-Fidelity Accelerated Design of High-performance Electrochemical Systems

Large-scale electrification is vital to addressing the climate crisis, but several scientific and technological challenges remain to fully electrify both the chemical industry and transportation. In both of these areas, new electrochemical materials will be critical, but their development currently relies heavily on human-time-intensive experimental trial and error and computationally expensive first-principles, meso-scale and continuum simulations. To accelerate this process, our team has developed the AutoMat platform. AutoMat can accelerate development of new electrochemical materials along two avenues: first, automated input generation and management of simulations at multiple lengthscales as well as “handoff” of outputs from one lengthscale as inputs to the next; and second, replacement of the most computationally intensive simulation processes with machine-learned surrogate models. The crux of our team’s effort was not “reinventing the wheel” by developing entirely new techniques, but rather building a “superhighway” that allows existing state-of-the-art techniques to run faster and more smoothly than before. AutoMat can utilize tools spanning from first-principles quantum chemistry computations to automated robotic experimentation, and is driven by design space search techniques to reduce the number of iterations through the full simulation loop by rapidly targeting promising regions of design spaces such as single-atom alloy catalysts or blends of liquid electrolytes.

25 ENERGY STORAGE↗

Numerical Investigation of Fluid Flow and Space Charge in Liquid Argon Time Projection Chamber (LArTPC) Detectors

Overview This project focused on developing a high-fidelity numerical framework to simulate the multiphysics environment within Liquid Argon Time Projection Chamber (LArTPC) detectors. The primary objective was to characterize the complex interplay between ion transport, background fluid dynamics, and electric field distortions—a critical factor for the calibration and sensitivity of next-generation High Energy Physics experiments, such as DUNE. Technical Achievements The research successfully yielded a hybrid numerical space-charge solver utilizing a Cell-Centered Finite Volume Method (FVM) for ion transport coupled with a Finite Element Method (FEM) for electric potential. Key accomplishments include: • Verification & Validation: The 3-D solver was rigorously verified against 1-D analytical solutions, demonstrating high numerical accuracy in predicting space-charge-induced field deviations. • Field Distortion Analysis: 3D simulations revealed that space charge effects introduce significant non-uniformities in the electric field. Critically, the research identified that background LAr flow velocities, when comparable to ion drift velocities, markedly exacerbate these distortions. • Technology Transfer: The resulting source code and comprehensive user manuals were successfully transferred to collaborators at Fermilab, providing a portable computational tool for the broader scientific community. Challenges and Future Directions While the space-charge solver achieved all performance metrics, the integrated fluid dynamics modeling encountered convergence challenges stemming from the extreme 200-fold disparity in length scales between the detector's 37 mm inlet pipes and the 8-meter global domain. To address this, the project has identified a clear technical pivot toward Hierarchical Geometric Adaptive Mesh Refinement (HG-AMR). By implementing an h-type refinement strategy with hanging nodes, future iterations of this solver will be capable of resolving localized high-gradient inlet flows without the prohibitive computational costs of regular grids. This advancement, combined with data-driven uncertainty quantification based on MicroBooNE-style calibration, will enable the precise modeling of detector responses in large-scale cryogenic environments where direct measurement remains difficult. Impact The computational tools developed under this award provide a foundation for enhancing the energy resolution and spatial reconstruction of noble liquid detectors. By bridging the gap between theoretical fluid dynamics and experimental field calibration, this work supports the DOE’s mission to advance the frontiers of neutrino physics and dark matter detection.

42 ENGINEERING↗

Machine Learning-Driven Conservative-to-Primitive Conversion in Hybrid Piecewise Polytropic and Tabulated Equations of State

We present a novel machine learning (ML)-based method to accelerate conservative-to-primitive inversion, focusing on hybrid piecewise polytropic and tabulated equations of state. Traditional root-finding techniques are computationally expensive, particularly for large-scale relativistic hydrodynamics simulations. To address this, we employ feedforward neural networks (NNC2PS and NNC2PL), trained in PyTorch (2.0+) and optimized for GPU inference using NVIDIA TensorRT (8.4.1), achieving significant speedups with minimal accuracy loss. The NNC2PS model achieves 𝐿 1 and 𝐿 ∞ errors of 4.54 × 10 −7 and 3.44 × 10−6, respectively, while the NNC2PL model exhibits even lower error values. TensorRT optimization with mixed-precision deployment substantially accelerates performance compared to traditional root-finding methods. Specifically, the mixed-precision TensorRT engine for NNC2PS achieves inference speeds approximately 400 times faster than a traditional single-threaded CPU implementation for a dataset size of 1,000,000 points. Ideal parallelization across an entire compute node in the Delta supercomputer (dual AMD 64-core 2.45 GHz Milan processors and 8 NVIDIA A100 GPUs with 40 GB HBM2 RAM and NVLink) predicts a 25-fold speedup for TensorRT over an optimally parallelized numerical method when processing 8 million data points. Moreover, the ML method exhibits sub-linear scaling with increasing dataset sizes. We release the scientific software developed, enabling further validation and extension of our findings. By exploiting the underlying symmetries within the equation of state, these findings highlight the potential of ML, combined with GPU optimization and model quantization, to accelerate conservative-to-primitive inversion in relativistic hydrodynamics simulations.

conservative-to-primitive conversion↗

Local isotropy in distorted turbulent boundary layers at high Reynolds number

This is a report on the continuation of our experimental investigations of the hypothesis of local isotropy in shear flows. This hypothesis, which states that at sufficiently high Reynolds numbers the small-scale structures of turbulent motions are independent of large-scale structures and mean deformations, has been used in theoretical studies of turbulence and computational methods such as large-eddy simulation. Since Kolmogorov proposed his theory, there have been many experiments, conducted in wakes, jets, mixing layers, a tidal channel, and atmospheric and laboratory boundary layers, in which attempts have been made to verify - or refute - the local-isotropy hypothesis. However, a review of the literature over the last five decades indicated that, despite all these experiments in shear flows, there was no consensus in the scientific community regarding this hypothesis, and, therefore, it seemed worthwhile to undertake a fresh experimental investigation into this question.

Saddoughi, Seyed G.↗

Rationalizing Burned Carbon with Carbon Monoxide Exported from South America

We present several estimates cross-checking the fluxes of carbon to the atmosphere from burning, comparing models that are based on simple land-surface parameterizations and atmospheric transport dynamics. Both estimates made by NASA Ames and USP modeling techniques are quite high compared to some detailed satellite/land-use studies of emissions. The flux of carbon liberated to the atmosphere via biomass burning is important for several reasons. This flux is a fundamental statistic for the parameterization of the large-scale flux of gases controlling the reactive greenhouse gases methane and ozone. Similarly, it is central to the estimation of the translocation of nitrogen and pyrodenitrification in the tropics. Thirdly, CO2 emitted from rainforest clearing contributes directly to carbon lost from the rainforest system as it contributes to greenhouse gas forcing. While CO2 from pasturage, agriculture, etc, is considered to be reabsorbed seasonally, and so "off budget" for the carbon cycle, it must also be accounted. CO2 anomalies related to daily weather and interannual climatic variation are strong enough to perturb our scientific perception of long-term carbon storage trends. We compare fluxes deduced from land-use statistics (originally, W.M. Hao) and from satellite hot pixels (A. Setzer) with atmospheric fluxes determined by the mesoscale/continental scale models RAMS and MM5, and point to some new work with highly resolved global models (the NASA Data Assimilation Office's GEOS4). Our simulations are tied to events, so that measured tracers like CO tie the models directly to the burning and meteorology of a specific period. We point out a particular sensitivity in estimates based on CO, and indicate how analysis of CO2 along with other biomass-burning tracers may lead to an improved multi-species estimator of carbon burned.

Chatfield, R.↗

Graphics Processing Unit (GPU) Acceleration of the Goddard Earth Observing System Atmospheric Model

The Goddard Earth Observing System 5 (GEOS-5) is the atmospheric model used by the Global Modeling and Assimilation Office (GMAO) for a variety of applications, from long-term climate prediction at relatively coarse resolution, to data assimilation and numerical weather prediction, to very high-resolution cloud-resolving simulations. GEOS-5 is being ported to a graphics processing unit (GPU) cluster at the NASA Center for Climate Simulation (NCCS). By utilizing GPU co-processor technology, we expect to increase the throughput of GEOS-5 by at least an order of magnitude, and accelerate the process of scientific exploration across all scales of global modeling, including: The large-scale, high-end application of non-hydrostatic, global, cloud-resolving modeling at 10- to I-kilometer (km) global resolutions Intermediate-resolution seasonal climate and weather prediction at 50- to 25-km on small clusters of GPUs Long-range, coarse-resolution climate modeling, enabled on a small box of GPUs for the individual researcher After being ported to the GPU cluster, the primary physics components and the dynamical core of GEOS-5 have demonstrated a potential speedup of 15-40 times over conventional processor cores. Performance improvements of this magnitude reduce the required scalability of 1-km, global, cloud-resolving models from an unfathomable 6 million cores to an attainable 200,000 GPU-enabled cores.

Putnam, Williama↗

The XXL Survey I. Scientific Motivations - Xmm-Newton Observing Plan - Follow-up Observations and Simulation Programme

The quest for the cosmological parameters that describe our universe continues to motivate the scientific community to undertake very large survey initiatives across the electromagnetic spectrum. Over the past two decades, the Chandra and XMM-Newton observatories have supported numerous studies of X-ray-selected clusters of galaxies, active galactic nuclei (AGNs), and the X-ray background. The present paper is the first in a series reporting results of the XXL-XMM survey; it comes at a time when the Planck mission results are being finalized. Aims. We present the XXL Survey, the largest XMM programme totaling some 6.9 Ms to date and involving an international consortium of roughly 100 members. The XXL Survey covers two extragalactic areas of 25 deg2 each at a point-source sensitivity of approx. 5 x 10(exp 15) erg/s/sq cm in the [0.5-2] keV band (completeness limit). The surveys main goals are to provide constraints on the dark energy equation of state from the space-time-distribution of clusters of galaxies and to serve as a pathfinder for future, wide-area X-ray missions. We review science objectives, including cluster studies, AGN evolution, and large-scale structure, that are being conducted with the support of approximately 30 follow-up programs. Methods. We describe the 542 XMM observations along with the associated multi- and numerical simulation programmes. We give a detailed account of the X-ray processing steps and describe innovative tools being developed for the cosmological analysis. Results. The paper provides a thorough evaluation of the X-ray data, including quality controls, photon statistics, exposure and background maps, and sky coverage. Source catalogue construction and multi-associations are briefly described. This material will be the basis for the calculation of the cluster and AGN selection functions, critical elements of the cosmological and science analyses. Conclusions. The XXL multi- data set will have a unique lasting legacy value for cosmological and extragalactic studies and will serve asa calibration resource for future dark energy studies with clusters and other X-ray selected sources. With the present article, we release the XMM XXL photon and smoothed images along with the corresponding exposure maps.

X-rays: general – large-scale structure of Unive↗