Engineering PapersSearch

DOE OSTI · 3368791

Scalable Hybrid Learning Techniques for Scientific Data Compression

Abstract

Data compression is becoming critical for storing scientific data because many scientific applications need to store large amounts of data and post process this data for scientific discovery. Unlike image and video compression algorithms that limit errors to primary data (PD), scientists require compression techniques that accurately preserve derived quantities of interest (QoIs). Here, this article presents a physics-informed compression technique implemented as an end-to-end, scalable, GPU-based pipeline for data compression that addresses this requirement. Our hybrid compression technique combines machine learning techniques and standard compression methods. Specifically, we combine an autoencoder, an error-bounded lossy compressor to provide guarantees on raw data error, and a constraint satisfaction post-processing step to preserve the QoIs within a minimal error (generally less than floating point error). The effectiveness of the data compression pipeline is demonstrated by compressing nuclear fusion simulation data generated by a large-scale fusion code, XGC, which produces hundreds of terabytes of data in a single day. Our approach works within the ADIOS framework and results in compression by a factor of more than 150 while requiring only a few percent of the computational resources necessary for generating the data, making the overall approach highly effective for practical scenarios.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Banerjee, Tania [Univ. of Florida, Gainesville, FL (United States)] (ORCID:0000000347370001), Choi, Jong Youl [Oak Ridge National Laboratory (ORNL), Oak Ridge, TN (United States)] (ORCID:0000000264596152), Lee, Jaemoon [Univ. of Florida, Gainesville, FL (United States)] (ORCID:0000000298689410), Gong, Qian [Oak Ridge National Laboratory (ORNL), Oak Ridge, TN (United States)] (ORCID:0000000235704142), Chen, Jieyang [Oak Ridge National Laboratory (ORNL), Oak Ridge, TN (United States)] (ORCID:0000000219059171), Klasky, Scott [Oak Ridge National Laboratory (ORNL), Oak Ridge, TN (United States)] (ORCID:0000000335595772), Rangarajan, Anand [Univ. of Florida, Gainesville, FL (United States)] (ORCID:0000000186958436), Ranka, Sanjay [Univ. of Florida, Gainesville, FL (United States)] (ORCID:0000000348861988). 2025-10-28. Scalable Hybrid Learning Techniques for Scientific Data Compression. https://doi.org/10.1109/tpds.2025.3623935

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related reports

Drift-kinetic effects of tungsten on plasma response to RMP in ITER

Here, effects of high- Z ( Z is the particle charge number) tungsten impurity ions on the plasma response to the resonant magnetic perturbation (RMP) field are numerically investigated for the ITER 15 MA baseline scenario, where the tungsten contribution to the plasma response is computed with a drift-kinetic model while the bulk thermal particle contributions follow the fluid approximation. The study yields three highlights: (i) the drift-kinetic contribution of the tungsten impurity exerts minor influence on the plasma response compared to that computed by the pure fluid model without tungsten; (ii) a new figure of merit, based on the resonant spectrum perturbation at the plasma boundary surface, results in different optimal coil phasing compared to that previously obtained by maximizing the edge-peeling plasma response; (iii) the optimal $n = 3$ RMP (for edge localized mode (ELM) control, $n$ is the toroidal mode number) is found to induce a large tungsten particle influx near the plasma edge associated with the neoclassical toroidal viscosity. The study thus provides useful data on the compatibility of the full tungsten wall with RMP ELM control in ITER.

ITER

Modelling the limiter ramp-up of WEST for addressing the future challenges of ITER

This paper presents a joint experimental and numerical investigation into the physics of long limited plasma ramp-up in tokamaks with tungsten (W) first walls, a critical phase for ITER operation. The comparison between the average plasma quantities simulated using the SolEdge-HDG code and the measurements taken during three successive WEST discharges after boronisation shows how challenging it is to predict this phase. While simulations reproduce the general trends at the midplane, with reasonable match in density profiles, they consistently underestimate core temperatures possibly due to too large perpendicular heat conductivity. On the contrary, at the high field side (HFS) limiter, simulations overestimate the measured quantities, and highlights the limitation of using Bohm boundary conditions at grazing magnetic angles. Experimental measurements reveal that the boron layer is rapidly eroded, on a timescale comparable to a single ITER discharge. The subsequent transition from a boron-coated to a tungsten wall increases recycling and significantly degrades the core plasma, reducing the electron temperature by nearly half due to W contamination, despite wall parameters remaining stable. Furthermore, comparisons with Langmuir probes, bolometry, reflectometry, and spectroscopy indicate that the experimental far scrape-Off layer (SOL) is significantly wider than simulated. This wide SOL implies that boron erosion extends along the entire HFS limiter rather than being confined to the contact point. This work highlights some characteristics of the plasma during this phase of the discharge and emphasizes the current modelling issues that need to be resolved in order to obtain reliable predictions concerning the ITER ramp-up.

ITER

Predicting core transport in ITER baseline discharges with neon injections

Achieving self-consistent performance predictions for ITER requires integrated modeling of core transport and divertor power exhaust under realistic impurity conditions. We present results from a systematic power-flow and impurity-content study for the ITER 15 MA baseline scenario constrained directly by existing SOLPS-ITER neon-seeded divertor solutions. Using the OMFIT STEP workflow, stationary temperature and density profiles are predicted with TGYRO for $1.5 \unicode{x2A7D} Z_\textrm{eff} \unicode{x2A7D} 2.5$, and the corresponding power crossing the separatrix $P_\textrm{sep}$ is evaluated. We find that $P_\textrm{sep}$ varies by more than a factor of 1.7 across this scan and matches the ${\sim}100$ MW SOLPS-ITER prediction when $Z_\textrm{eff} \simeq 1.6$ or when auxiliary heating is reduced to ${\sim}75\%$ of nominal. Rotation-sensitivity studies show that plausible variations in toroidal flow magnitude modify $P_\textrm{sep}$ by $\lesssim 20\%$, while AURORA modeling confirms that charge-exchange radiation inside the separatrix is dynamically negligible under predicted ITER neutral densities. These results identify a restricted compatibility window, $Z_\textrm{eff} \approx 1.6$ –1.75 and $0.75 \lesssim f_{P_\textrm{aux}} \unicode{x2A7D} 1.0$, in which core transport predictions remain aligned with neon-seeded divertor protection targets. This self-consistent, model-constrained framework provides actionable guidance for impurity control and auxiliary-heating scheduling in early ITER operation and supports future whole-device scenario optimization.

ITER