Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “aurora”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Multi-task Parallelism for Robust Pre-training of Graph Foundation Models on Multi-source, Multi-fidelity Atomistic Modeling Data

Graph foundation models using graph neural networks promise sustainable, efficient atomistic modeling. To tackle challenges of processing multi-source, multi-fidelity data during pre-training, recent studies employ multi-task learning, in which shared message passing layers initially process input atomistic structures regardless of source, then route them to multiple decoding heads that predict data-specific outputs. This approach stabilizes pre-training and enhances a model’s transferability to unexplored chemical regions. Preliminary results on approximately four million structures are encouraging, yet questions remain about generalizability to larger, more diverse datasets and scalability on supercomputers. We propose a multi-task parallelism method that distributes each head across computing resources with GPU acceleration. Implemented in the open-source HydraGNN architecture, our method was trained on over 24 million structures from five datasets and tested on the Perlmutter, Aurora, and Frontier supercomputers, demonstrating efficient scaling on all three highly heterogeneous super-computing architectures.

Lupo Pasini, Massimiliano [ORNL] (ORCID:0000000249↗

SoK: What does it Mean to Benchmark Database Forensics?

Relational Database Management Systems are the backbone of modern enterprises and public-sector services, and are thus frequent targets of security incidents, insider threats, and thorough regulatory audits. Consequently, databases have become key sources of digital evidence, requiring investigators to reconstruct past activity from audit logs, transaction logs, and backups. Although benchmarking frameworks such as those developed by the Transaction Processing Performance Council (TPC) are widely used to evaluate database performance, they do not capture forensic requirements such as evidentiary completeness, tamper-evidence, chain of custody, or regulatory compliance under GDPR and CCPA. This survey examines the emerging domain of forensic database benchmarking. We gathered prior research on database forensics, secure logging, and tamper-evident data structures; we analyze modern forensic-ready features in commercial and open-source systems (SQL Server Ledger, Oracle Blockchain Tables, PostgreSQL pgAudit, Db2 Audit, Aurora Database Activity Streams, Oracle Real Application Security and IBM Guardium) and assess why existing benchmarks are insufficient. We propose forensic workloads, metrics, and methodologies that incorporate adversarial stressors, deleted-record recovery, and backup analysis. We also identify open research problems and call for a community-driven forensic benchmark suite. The result is an idea for evaluating not only database performance but also forensic soundness, bridging the gap between system engineering, compliance, and digital investigations.

Lenard, Ben↗

Modelling of the electron cyclotron emission burst from a laboratory tokamak plasma with loss-cone maser instability

The maser instability associated with the loss-cone distribution has been widely invoked to explain the radio bursts observed in the astrophysical plasma environment, such as aurora and corona. In the laboratory plasma of a tokamak, events reminiscent of these radio bursts have also been frequently observed as an electron cyclotron emission (ECE) burst in the microwave range (~2f ce near the last closed flux surface) during transient magnetohydrodynamic events. These bursts have a short duration of ~10 μs and display a radiation spectrum corresponding to a radiation temperature T e,rad of over 30 keV while the edge thermal electron temperature T e is only in the range of 1 keV. Suprathermal electrons can be generated through magnetic reconnection, and a loss-cone distribution can be generated through open stochastic field lines in the magnetic mirror of the near-edge region of a tokamak plasma. Radiation modelling shows that a sharp distribution gradient ∂f/∂v ⊥ > 0 at the loss-cone boundary can cause a negative absorption of ECE radiation through the maser instability. The negative absorption then amplifies the radiation so that the microwave intensity is significantly stronger than the thermal value. The significant T e,rad from the simulations suggests the potential role of the loss-cone maser instability in generating the ECE burst in a tokamak.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Enabling Multireference Calculations on Multimetallic Systems with Graphic Processing Units

Modeling multimetallic systems efficiently enables faster prediction of desirable chemical properties and the design of new materials. This work describes an initial implementation for performing multireference wave function method localized active-space self-consistent field (LASSCF) calculations through the use of multiple graphics processing units (GPUs) to accelerate time-to-solution. Density fitting is leveraged to reduce memory requirements, and we demonstrate the ability to fully utilize multi-GPU compute nodes. Performance improvements of 5–10x in total application runtime were observed in LASSCF calculations for multimetallic catalyst systems up to 1200 AOs and an active space of (22e,40o) using up to four NVIDIA A100 GPUs. Furthermore, written with performance portability in mind, a comparable performance is also observed in early runs on the Aurora exascale system using Intel Max Series GPUs.

Algorithms↗

62°S Witnesses the Transition of Boundary Layer Marine Aerosol Pattern Over the Southern Ocean (50°S–68°S, 63°E– 150°E) During the Spring and Summer: Results From MARCUS (I)

The Atmospheric Radiation Measurement Mobile Facility-2 was installed onboard the research vessel Aurora Australis to measure aerosol properties during the 2017-2018 Measurement of Aerosols, Radiation, and CloUds over the pristine Southern ocean (MARCUS) Experiment, providing unique data on aerosols latitudinal and seasonal variation, including south of 60 degrees S where previous observations are scarce. Data from a Cloud Condensation Nuclei (CCN) counter and Ultra-High-Sensitivity Aerosol Spectrometer show that both the number concentration (N-CCN) and size distribution of CCN-active aerosols, with diameters (D) between 60 nm < D < 1,000 nm are different over the North Southern Ocean (NSO) (50 degrees S-60 degrees S) and the South Southern Ocean (SSO) (62 degrees S-68 degrees S). The average NSO N-CCN at 0.2% and 0.5% supersaturation were 28% and 49% less than that over the SSO. This increase of CCN over the SSO is caused by the increase of aerosols with 60 nm < D < 200 nm, consistent with calculations of Aerosol Scattering Angstrom Exponents derived from a nephelometer. Aerosol hygroscopicity growth factor measured by the Hygroscopic Tandem Differential Mobility Analyzer stayed close to 1.41 for aerosols with 50 nm < D < 250 nm over the SSO, but increased from 1.30 to 1.67 over the NSO, indicating different chemical compositions. Both CCN and Ice Nucleating Particles (INPs) showed a stronger variation with season than with latitude. The variation of heat-labile and presumably proteinacous INPs suggests an increase of ice nucleating-active microbes in summer.

Niu, Qing↗

Giant Undulations Driven by Pitch‐Angle Scattering of Time Domain Structures Modulated by Plasmapause Surface Wave

Abstract Plasmapause surface waves (PSWs) near the plasmapause boundary are regarded to be the magnetospheric source of ionospheric auroral giant undulations (GUs) located at the equatorward boundary of diffuse aurora. However, the observational evidence of wave‐particle interaction connecting PSWs and GUs is absent. In this letter, we demonstrate GUs are driven by pitch‐angle scattering of time domain structures modulated by the PSWs, based on the conjugated ionospheric and magnetospheric observations. Specifically, ionospheric GUs are lighted by the pitch‐angle scattering of <1 keV thermal electron and ions and energetic ions with energy up to dozens of keV near the plasmapause. Further, the total fluxes during one PSW period and energy of scattered electron and ions determine the size and luminosity of GUs. Our research provides observational evidence that PSWs cause periodic electron precipitation via modulating the time domain structures rather than the previously predicted chorus or electron cyclotron harmonic waves.

Zhou, Yi‐Jia [Weihai Institute for Interdisciplina↗

Auroral and Non‐Auroral H 3 + Ion Winds at Uranus With Keck‐NIRSPEC and IRTF‐iSHELL

Abstract To date, no investigation has documented ionospheric flows at Uranus. Previous investigations of Jupiter and Saturn have demonstrated that mapping ion winds can be used to understand ionospheric currents and how these connect to magnetosphere‐ionosphere coupling. We present a study of Uranus's near infrared emissions (NIR) using data from the Keck II Telescope's Near InfraRed SPECtrograph (NIRSPEC) and the InfraRed Telescope Facility's iSHELL spectrograph. H 3 + emission lines were used to derive dawn‐to‐dusk intensity, ionospheric temperatures and ion densities to identify auroral emissions, with their Doppler shifts used to measure ion velocities. We confirm the presence of the southern NIR aurora in 2016, driven by elevated H 3 + column densities up to 6.0 × 10 16 m −2 . While no auroral emissions were detected in 2014, we find a 14%–20% super rotation across the planet's disk in 2014 and a 7%–18% super rotation in 2016.

Thomas, Emma M. [Department of Mathematics Physics↗

Characterization and controllability of radiated power via extrinsic impurity seeding in strongly negative triangularity plasmas in DIII-D

Experiments with extrinsic impurity seeding in strongly negative triangularity shapes in DIII-D achieved radiated power fractions (relative to input power) of up to ≈85% total radiation and ≈55% core radiation in steady-operating conditions. The relationship between core and total radiation was sensitive to impurity species and input power. Attempts to reach higher radiation levels via higher impurity flows resulted in radiative collapse disruptions. Nitrogen, neon, argon, and krypton were tested. Injection was by gas puffing, usually controlled by feeding back real-time estimates for core or total radiating fractions ($P_\textrm{rad} / P_\textrm{input}$). Argon and krypton were controlled by feeding back total radiating fraction, whereas neon was controlled by feeding back the core radiating fraction, due to the lower efficiency of neon as a divertor radiator. Nitrogen flows were pre-programmed. Reasonable total $P_\textrm{rad}$ control target following was achieved with argon or krypton. Control was more challenging with the neon/core radiation configuration, which was more prone to slow response and overshooting of the control target. Poor particle removal contributed to the control challenge: neon particle inventory within the last closed flux surface was roughly constant for up to 1 s (the longest duration tested) after neon injection was halted. However, separate experiments with laser-blow off of non-recycling impurities measured a short impurity confinement time, on the scale of the energy confinement time of ∼100 ms. Modeling with the Aurora impurity transport simulation matched experimental neon density profiles with full recycling (R = 1.0) and weak pumping but predicted rapid decreases in neon inventory if pumping were increased or recycling decreased. This indicates that changes outside the confined plasma (adding a well-placed pump) would improve controllability for all highly recycling species.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Predicting core transport in ITER baseline discharges with neon injections

Achieving self-consistent performance predictions for ITER requires integrated modeling of core transport and divertor power exhaust under realistic impurity conditions. We present results from a systematic power-flow and impurity-content study for the ITER 15 MA baseline scenario constrained directly by existing SOLPS-ITER neon-seeded divertor solutions. Using the OMFIT STEP workflow, stationary temperature and density profiles are predicted with TGYRO for $1.5 \unicode{x2A7D} Z_\textrm{eff} \unicode{x2A7D} 2.5$, and the corresponding power crossing the separatrix $P_\textrm{sep}$ is evaluated. We find that $P_\textrm{sep}$ varies by more than a factor of 1.7 across this scan and matches the ${\sim}100$ MW SOLPS-ITER prediction when $Z_\textrm{eff} \simeq 1.6$ or when auxiliary heating is reduced to ${\sim}75\%$ of nominal. Rotation-sensitivity studies show that plausible variations in toroidal flow magnitude modify $P_\textrm{sep}$ by $\lesssim 20\%$, while AURORA modeling confirms that charge-exchange radiation inside the separatrix is dynamically negligible under predicted ITER neutral densities. These results identify a restricted compatibility window, $Z_\textrm{eff} \approx 1.6$ –1.75 and $0.75 \lesssim f_{P_\textrm{aux}} \unicode{x2A7D} 1.0$, in which core transport predictions remain aligned with neon-seeded divertor protection targets. This self-consistent, model-constrained framework provides actionable guidance for impurity control and auxiliary-heating scheduling in early ITER operation and supports future whole-device scenario optimization.

ITER↗

Search for Dark Matter Ionization on the Night Side of Jupiter with Cassini

We present a new search for dark matter (DM) using planetary atmospheres. We point out that annihilating DM in planets can produce ionizing radiation, which can lead to excess production of ionospheric H 3 + . We apply this search strategy to the night side of Jupiter near the equator. The night side has zero solar irradiation, and low latitudes are sufficiently far from ionizing auroras, leading to a low-background search. We use data on ionospheric H 3 + emission collected three hours either side of Jovian midnight, during its flyby in 2000, and set novel constraints on the DM-nucleon scattering cross section down to about 10 − 38 cm 2 . We also highlight that DM atmospheric ionization may be detected in Jovian exoplanets using future high-precision measurements of planetary spectra. Published by the American Physical Society 2024

79 ASTRONOMY AND ASTROPHYSICS↗

In-Transit Data Transport Strategies for Coupled AI-Simulation Workflow Patterns

Coupled AI-Simulation workflows are becoming the major workloads for HPC facilities, and their increasing complexity necessitates new tools for performance analysis and prototyping of new in-situ workflows. We present SimAI-Bench, a tool designed to both prototype and evaluate these coupled workflows. In this paper, we use SimAI-Bench to benchmark the data transport performance of two common patterns on the Aurora supercomputer: a one-to-one workflow with co-located simulation and AI training instances, and a many-to-one workflow where a single AI model is trained from an ensemble of simulations. For the one-to-one pattern, our analysis shows that node-local and DragonHPC data staging strategies provide excellent performance compared Redis and Lustre file system. For the many-to-one pattern, we find that data transport becomes a dominant bottleneck as the ensemble size grows. Our evaluation reveals that file system is the optimal solution among the tested strategies for the many-to-one pattern.

Tummalapalli, Harikrishna [Argonne National Labora↗

HydraGNN v4.0

The new version of HydraGNN v4.0 provides additional core capabilities, such as: Inclusion of multi-body atomistic cluster expansion MACE, polarizable atom interaction neural network PAINN, and equivariant principal neighborhood aggregation (PNAEq) among the message passing layers supported -Inclusion of graph transformers to directly model long-range interactions between nodes that are distant in the graph topology Integration of graph transformers with message passing layers by combining the graph embedding generated by the two mechanisms, which allows for an improved expressivity of the HydraGNN architecture Improved re-implementation of multi-task learning (MTL) to allow its use for stabilized training across imbalanced, multi-source, multi-fidelity data Introduction of multi-task parallelism, a newly proposed type of model parallelism specifically for MTL architectures, which allows to dispatch different output decoding heads to different GPU devices Integration of multi-task parallelism with pre-existing distributed data parallelism to enable a 2D parallelization for distributed training Improved portability of the distributed training across Intel GPUs, which has been testes on ALCF exascale supercomputer Aurora Inclusion of 2-level fine-grained energy profilers portable across NVIDIA, AMD, and Intel GPUs to monitor the power and energy consumption associated with different functions executed by the HydraGNN code during data pre-load and training Restructuring of previous examples and inclusion of new sets of examples to illustrate the download, preprocess, and training of HydraGNN models on new large-scale open-source datasets for atomistic materials modeling (e.g., Alexandria, Transition1x, OMat24, OMol25)

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

HydraGNN v5.0

HydraGNN v5.0 expands the code base into a more portable, scalable, and flexible framework for scientific graph learning, with particular strength in atomistic machine-learning interatomic potentials and large-scale distributed training. The release adds Fully Sharded Data Parallel (FSDP) support alongside existing DDP and DeepSpeed paths, including FSDP-aware checkpointing and optimizer integration, and introduces a configurable multi-precision training workflow supporting FP32, BF16, and FP64 across GPUs and Intel XPUs. For atomistic modeling, HydraGNN v5.0 strengthens its MLIP capabilities through dynamic graph construction at every forward pass, energy-conserving force prediction via automatic differentiation, and per-atom energy loss formulations, while extending EGNN models to properly handle periodic boundary conditions. The release also broadens model expressiveness through graph-level attribute conditioning, adds new multi-task and model-parallel extensions such as MACE support and encoder/decoder branch optimization, and expands application coverage with integrated examples for datasets including OC25, Nabla2-DFT, QCML, Open Polymers 2026, and OPF. In parallel, HydraGNN v5.0 improves production readiness through performance optimizations for large-scale runs, stratified sampling and linear-regression preprocessing utilities, and tested installation scripts for DOE supercomputers including Frontier, Aurora, Perlmutter, and Andes. Overall, the release advances HydraGNN as a robust software platform for scalable graph neural networks across materials science, chemistry, and scientific machine learning workflows

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

A graphics processing unit accelerated sparse direct solver and preconditioner with block low rank compression

We present the GPU implementation efforts and challenges of the sparse solver package STRUMPACK. The code is made publicly available on github with a permissive BSD license. STRUMPACK implements an approximate multifrontal solver, a sparse LU factorization which makes use of compression methods to accelerate time to solution and reduce memory usage. Multiple compression schemes based on rank-structured and hierarchical matrix approximations are supported, including hierarchically semi-separable, hierarchically off-diagonal butterfly, and block low rank. Here, in this paper, we present the GPU implementation of the block low rank (BLR) compression method within a multifrontal solver. Our GPU implementation relies on highly optimized vendor libraries such as cuBLAS and cuSOLVER for NVIDIA GPUs, rocBLAS and rocSOLVER for AMD GPUs and the Intel oneAPI Math Kernel Library (oneMKL) for Intel GPUs. Additionally, we rely on external open source libraries such as SLATE (Software for Linear Algebra Targeting Exascale), MAGMA (Matrix Algebra on GPU and Multi-core Architectures), and KBLAS (KAUST BLAS). SLATE is used as a GPU-capable ScaLAPACK replacement. From MAGMA we use variable sized batched dense linear algebra operations such as GEMM, TRSM and LU with partial pivoting. KBLAS provides efficient (batched) low rank matrix compression for NVIDIA GPUs using an adaptive randomized sampling scheme. The resulting sparse solver and preconditioner runs on NVIDIA, AMD and Intel GPUs. Interfaces are available from PETSc, Trilinos and MFEM, or the solver can be used directly in user code. We report results for a range of benchmark applications, using the Perlmutter system from NERSC, Frontier from ORNL, and Aurora from ALCF. For a high frequency wave equation on a regular mesh, using 32 Perlmutter compute nodes, the factorization phase of the exact GPU solver is about 6.5× faster compared to the CPU-only solver. The BLR-enabled GPU solver is about 13.8× faster than the CPU exact solver. For a collection of SuiteSparse matrices, the STRUMPACK exact factorization on a single GPU is on average 1.9× faster than NVIDIA’s cuDSS solver.

97 MATHEMATICS AND COMPUTING↗

Portable, heterogeneous ensemble workflows at scale using libEnsemble

libEnsemble is a Python-based toolkit for running dynamic ensembles, developed as part of the DOE Exascale Computing Project. The toolkit utilizes a unique generator–simulator–allocator paradigm, where generators produce input for simulators, simulators evaluate those inputs, and allocators decide whether and when a simulator or generator should be called. The generator steers the ensemble based on simulation results. Generators may, for example, apply methods for numerical optimization, machine learning, or statistical calibration. libEnsemble communicates between a manager and workers. Flexibility is provided through multiple manager–worker communication substrates each of which has different benefits. These include Python’s multiprocessing, mpi4py, and TCP. Multisite ensembles are supported using Balsam or Globus Compute. We overview the unique characteristics of libEnsemble as well as current and potential interoperability with other packages in the workflow ecosystem. We highlight libEnsemble’s dynamic resource features: libEnsemble can detect system resources, such as available nodes, cores, and GPUs, and assign these in a portable way. These features allow users to specify the number of processors and GPUs required for each simulation; and resources will be automatically assigned on a wide range of systems, including Frontier, Aurora, and Perlmutter. Such ensembles can include multiple simulation types, some using GPUs and others using only CPUs, sharing nodes for maximum efficiency. We also describe the benefits of libEnsemble’s generator–simulator coupling, which easily exposes to the user the ability to cancel, and portably kill, running simulations based on models that are updated with intermediate simulation output. We demonstrate libEnsemble’s capabilities, scalability, and scientific impact via a Gaussian process surrogate training problem for the longitudinal density profile at the exit of a plasma accelerator stage. In conclusion, the study uses gpCAM for the surrogate model and employs either Wake-T or WarpX simulations, highlighting efficient use of resources that can easily extend to exascale.

Dynamic ensembles↗

PETSc/TAO developments for GPU-based early exascale systems

The Portable Extensible Toolkit for Scientific Computation (PETSc) library provides scalable solvers for nonlinear time-dependent differential and algebraic equations and for numerical optimization via the Toolkit for Advanced Optimization (TAO). PETSc is used in dozens of scientific fields and is an important building block for many simulation codes. During the U.S. Department of Energy’s Exascale Computing Project, the PETSc team has made substantial efforts to enable efficient utilization of the massive fine-grain parallelism present within exascale compute nodes and to enable performance portability across exascale architectures. We recap some of the challenges that designers of numerical libraries face in such an endeavor, and then discuss the many developments we have made, which include the addition of new GPU backends, features supporting efficient on-device matrix assembly, better support for asynchronicity and GPU kernel concurrency, and new communication infrastructure. In conclusion, we evaluate the performance of these developments on some pre-exascale systems as well as the early exascale systems Frontier and Aurora, using compute kernel, communication layer, solver, and mini-application benchmark studies, and then close with a few observations drawn from our experiences on the tension between portable performance and other goals of numerical libraries.

Exascale Computing Project (ECP)↗

Preparing MPICH for exascale

The advent of exascale supercomputers heralds a new era of scientific discovery, yet it introduces significant architectural challenges that must be overcome for MPI applications to fully exploit its potential. Among these challenges is the adoption of heterogeneous architectures, particularly the integration of GPUs to accelerate computation. Additionally, the complexity of multithreaded programming models has also become a critical factor in achieving performance at scale. The efficient utilization of hardware acceleration for communication, provided by modern NICs, is also essential for achieving low latency and high throughput communication in such complex systems. In response to these challenges, the MPICH library, a high-performance and widely used Message Passing Interface (MPI) implementation, has undergone significant enhancements. Here, this paper presents four major contributions that prepare MPICH for the exascale transition. First, we describe a lightweight communication stack that leverages the advanced features of modern NICs to maximize hardware acceleration. Second, our work showcases a highly scalable multithreaded communication model that addresses the complexities of concurrent environments. Third, we introduce GPU-aware communication capabilities that optimize data movement in GPU-integrated systems. Finally, we present a new datatype engine aimed at accelerating the use of MPI derived datatypes on GPUs. These improvements in the MPICH library not only address the immediate needs of exascale computing architectures but also set a foundation for exploiting future innovations in high-performance computing. By embracing these new designs and approaches, MPICH-derived libraries from HPE Cray and Intel were able to achieve real exascale performance on OLCF Frontier and ALCF Aurora respectively.

Guo, Yanfei [Argonne National Laboratory (ANL), Ar↗

Proactive Regulatory Approaches to Electrification and Load Growth: Workshop Report

On July 10 and 11, 2024, Pacific Northwest National Laboratory and RMI led a workshop in Aurora, Colorado, to explore novel and proactive approaches to electrification and load growth while minimizing risks and costs to customers. Over the next decade, a unique opportunity exists to invest strategically in the electricity system to enable electrification across the transportation, industrial, and building sectors and respond to data and technology-based load growth. However, current utility and regulatory planning practices are insufficient to identify and enable the right investments, and work must be done to reduce the risk and decisional uncertainty faced by utility regulatory commissions and utilities. Ensuring timely electrification investments may require new approaches to address risk, uncertainty, prudence, and cost recovery. Understanding the decision-making process and information needs of utilities and regulators is critical. New policies (or application of policies), financial tools, systems analysis, regulatory mechanisms, and enhanced process transparency may be required. The workshop's goal was to identify proactive regulatory approaches for electrification and load growth that minimize costs and risks to customers. Our intention was that the conversations and the resulting solutions and takeaways would be specific and tactical rather than general and theoretical and that together we would create actionable next steps for key actors in the system, including utilities, regulators, thought leaders, researchers, and the U.S. Department of Energy (DOE). This report is intended to provide workshop attendees with a record and summary of the discussion and proposals raised at the workshop and to provide interested entities who did not attend, such as other regulators, policymakers, utilities, and U.S. DOE offices, with an understanding of what was discussed and with ideas to explore in their organizations.

24 POWER TRANSMISSION AND DISTRIBUTION↗