Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Aurora”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Auroral and Non‐Auroral H 3 + Ion Winds at Uranus With Keck‐NIRSPEC and IRTF‐iSHELL

Abstract To date, no investigation has documented ionospheric flows at Uranus. Previous investigations of Jupiter and Saturn have demonstrated that mapping ion winds can be used to understand ionospheric currents and how these connect to magnetosphere‐ionosphere coupling. We present a study of Uranus's near infrared emissions (NIR) using data from the Keck II Telescope's Near InfraRed SPECtrograph (NIRSPEC) and the InfraRed Telescope Facility's iSHELL spectrograph. H 3 + emission lines were used to derive dawn‐to‐dusk intensity, ionospheric temperatures and ion densities to identify auroral emissions, with their Doppler shifts used to measure ion velocities. We confirm the presence of the southern NIR aurora in 2016, driven by elevated H 3 + column densities up to 6.0 × 10 16 m −2 . While no auroral emissions were detected in 2014, we find a 14%–20% super rotation across the planet's disk in 2014 and a 7%–18% super rotation in 2016.

Thomas, Emma M. [Department of Mathematics Physics↗

Simulations in the era of exascale computing

Exascale computers — supercomputers that can perform 10 18 floating point operations per second — started coming online in 2022: in the United States, Frontier launched as the first public exascale supercomputer and Aurora is due to open soon; OceanLight and Tianhe-3 are operational in China; and JUPITER is due to launch in 2023 in Europe. Supercomputers offer unprecedented opportunities for modelling complex materials. In this Viewpoint, five researchers working on different types of materials discuss the most promising directions in computational materials science.

36 MATERIALS SCIENCE↗

Creating and studying a scaled interplanetary coronal mass ejection

The Sun, being an active star, undergoes eruptions of magnetized plasma that reach the Earth and cause the aurorae near the poles. These eruptions, called coronal mass ejections (CMEs), send plasma and magnetic fields out into space. CMEs that reach planetary orbits are called interplanetary coronal mass ejections (ICMEs) and are a source of geomagnetic storms, which can cause major damage to our modern electrical systems with limited warning. To study ICME propagation, we devised a scaled experiment using the Big Red Ball (BRB) plasma containment device at the Wisconsin Plasma Physics Laboratory. These experiments inject a compact torus of plasma as an ICME through an ambient plasma inside the BRB, which acts as the interplanetary medium. Magnetic and temperature probes provide three-dimensional magnetic field information in time and space, as well as temperature and density as a function of time. Using this information, we can identify features in the compact torus that are consistent with those in real ICMEs. We also identify the shock, sheath, and ejecta similar to the structure of an ICME event. This experiment acts as a first step to providing information that can inform predictive models, which can give us time to shield our satellites and large electrical systems in the event that a powerful ICME were to strike.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A biophysical framework for double-drugging kinases

Selective orthosteric inhibition of kinases has been challenging due to the conserved active site architecture of kinases and emergence of resistance mutants. Simultaneous inhibition of distant orthosteric and allosteric sites, which we refer to as “double-drugging”, has recently been shown to be effective in overcoming drug resistance. However, detailed biophysical characterization of the cooperative nature between orthosteric and allosteric modulators has not been undertaken. Here, we provide a quantitative framework for double-drugging of kinases employing isothermal titration calorimetry, Förster resonance energy transfer, coupled-enzyme assays, and X-ray crystallography. We discern positive and negative cooperativity for Aurora A kinase (AurA) and Abelson kinase (Abl) with different combinations of orthosteric and allosteric modulators. We find that a conformational equilibrium shift is the main principle governing cooperativity. Notably, for both kinases, we find a synergistic decrease of the required orthosteric and allosteric drug dosages when used in combination to inhibit kinase activities to clinically relevant inhibition levels. X-ray crystal structures of the double-drugged kinase complexes reveal the molecular principles underlying the cooperative nature of double-drugging AurA and Abl with orthosteric and allosteric inhibitors. Finally, we observe a fully closed conformation of Abl when bound to a pair of positively cooperative orthosteric and allosteric modulators, shedding light on the puzzling abnormality of previously solved closed Abl structures. Collectively, our data provide mechanistic and structural insights into rational design and evaluation of double-drugging strategies.

59 BASIC BIOLOGICAL SCIENCES↗

Parallel interior-point solver for block-structured nonlinear programs on SIMD/GPU architectures

Here, we investigate how to port the standard interior-point method to new exascale architectures for block-structured nonlinear programs with state equations. Computationally, we decompose the interior-point algorithm into two successive operations: the evaluation of the derivatives and the solution of the associated Karush-Kuhn-Tucker (KKT) linear system. Our method accelerates both operations using two levels of parallelism. First, we distribute the computations on multiple processes using coarse parallelism. Second, each process uses SIMD/GPU accelerators locally to accelerate the operations using fine-grained parallelism. The KKT system is reduced by eliminating the inequalities and the state variables from the corresponding equations. We demonstrate our method's capability on the supercomputer Polaris, a testbed for the future exascale Aurora system. Each node is equipped with four GPUs, a setup amenable to our two-level approach. Our experiments on the stochastic optimal power flow problem show that the reduction method is 50x faster than the sparse linear solver HSL MA57 running in serial on the CPU, and 6x faster than Pardiso running in parallel on CPU on the same number of processes.

97 MATHEMATICS AND COMPUTING↗

Investigation of core impurity transport in DIII-D diverted negative triangularity plasmas

Abstract Tokamak operation at negative triangularity has been shown to offer high energy confinement without the typical disadvantages of edge pedestals (Marinoni et al 2021 Nucl. Fusion 61 116010). In this paper, we examine impurity transport in DIII-D diverted negative triangularity experiments. Analysis of charge exchange recombination spectroscopy reveals flat or hollow carbon density profiles in the core, and impurity confinement times consistently shorter than energy confinement times. Bayesian inferences of impurity transport coefficients based on laser blow-off injections and forward modeling via the Aurora package (Sciortino et al 2021 Plasma Phys. Control. Fusion 63 112001) show core cross-field diffusion to be higher in L-mode than in H-mode. Impurity profile shapes remain flat or hollow in all cases. Inferred radial profiles of diffusion and convection are compared to neoclassical, quasilinear gyrofluid, and nonlinear gyrokinetic simulations. Heat transport is observed to be better captured by reduced turbulence models with respect to particle transport. State-of-the-art gyrokinetic modeling compares favorably with measurements across multiple transport channels. Overall, these results suggest that diverted negative triangularity discharges may offer a path to a highly-radiative L-mode scenario with high core performance.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Characterization and controllability of radiated power via extrinsic impurity seeding in strongly negative triangularity plasmas in DIII-D

Experiments with extrinsic impurity seeding in strongly negative triangularity shapes in DIII-D achieved radiated power fractions (relative to input power) of up to ≈85% total radiation and ≈55% core radiation in steady-operating conditions. The relationship between core and total radiation was sensitive to impurity species and input power. Attempts to reach higher radiation levels via higher impurity flows resulted in radiative collapse disruptions. Nitrogen, neon, argon, and krypton were tested. Injection was by gas puffing, usually controlled by feeding back real-time estimates for core or total radiating fractions ($P_\textrm{rad} / P_\textrm{input}$). Argon and krypton were controlled by feeding back total radiating fraction, whereas neon was controlled by feeding back the core radiating fraction, due to the lower efficiency of neon as a divertor radiator. Nitrogen flows were pre-programmed. Reasonable total $P_\textrm{rad}$ control target following was achieved with argon or krypton. Control was more challenging with the neon/core radiation configuration, which was more prone to slow response and overshooting of the control target. Poor particle removal contributed to the control challenge: neon particle inventory within the last closed flux surface was roughly constant for up to 1 s (the longest duration tested) after neon injection was halted. However, separate experiments with laser-blow off of non-recycling impurities measured a short impurity confinement time, on the scale of the energy confinement time of ∼100 ms. Modeling with the Aurora impurity transport simulation matched experimental neon density profiles with full recycling (R = 1.0) and weak pumping but predicted rapid decreases in neon inventory if pumping were increased or recycling decreased. This indicates that changes outside the confined plasma (adding a well-placed pump) would improve controllability for all highly recycling species.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Predicting core transport in ITER baseline discharges with neon injections

Achieving self-consistent performance predictions for ITER requires integrated modeling of core transport and divertor power exhaust under realistic impurity conditions. We present results from a systematic power-flow and impurity-content study for the ITER 15 MA baseline scenario constrained directly by existing SOLPS-ITER neon-seeded divertor solutions. Using the OMFIT STEP workflow, stationary temperature and density profiles are predicted with TGYRO for $1.5 \unicode{x2A7D} Z_\textrm{eff} \unicode{x2A7D} 2.5$, and the corresponding power crossing the separatrix $P_\textrm{sep}$ is evaluated. We find that $P_\textrm{sep}$ varies by more than a factor of 1.7 across this scan and matches the ${\sim}100$ MW SOLPS-ITER prediction when $Z_\textrm{eff} \simeq 1.6$ or when auxiliary heating is reduced to ${\sim}75\%$ of nominal. Rotation-sensitivity studies show that plausible variations in toroidal flow magnitude modify $P_\textrm{sep}$ by $\lesssim 20\%$, while AURORA modeling confirms that charge-exchange radiation inside the separatrix is dynamically negligible under predicted ITER neutral densities. These results identify a restricted compatibility window, $Z_\textrm{eff} \approx 1.6$ –1.75 and $0.75 \lesssim f_{P_\textrm{aux}} \unicode{x2A7D} 1.0$, in which core transport predictions remain aligned with neon-seeded divertor protection targets. This self-consistent, model-constrained framework provides actionable guidance for impurity control and auxiliary-heating scheduling in early ITER operation and supports future whole-device scenario optimization.

ITER↗

Inference of main ion particle transport coefficients with experimentally constrained neutral ionization during edge localized mode recovery on DIII-D

Abstract The plasma and neutral density dynamics after an edge localized mode are investigated and utilized to infer the plasma transport coefficients for the density pedestal. The Lyman-Alpha Measurement Apparatus (LLAMA) diagnostic provides sub-millisecond profile measurements of the ionization and neutral density and shows significant poloidal asymmetries in both. Exploiting the absolute calibration of the LLAMA diagnostic allows quantitative comparison to the electron and main ion density profiles determined by charge-exchange recombination, Thomson scattering and interferometry. Separation of diffusion and convection contributions to the density pedestal transport are investigated through flux gradient methods and time-dependent forward modeling with Bayesian inference by adaptation of the Aurora transport code and IMPRAD framework to main ion particle transport. Both methods suggest time-dependent transport coefficients and are consistent with an inward particle pinch on the order of 1 m s −1 and diffusion coefficient of 0.05 m 2 s −1 in the steep density gradient region of the pedestal. While it is possible to recreate the experimentally observed phenomena with no pinch in the pedestal, low diffusion in the core and high outward convection in the near scrape-off layer are required without an inward pedestal pinch.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Pedestal main ion particle transport inference through gas puff modulation with experimental source measurements

Abstract Transport in the DIII-D high confinement mode (H-mode) pedestal is investigated through a periodic edge gas puff modulation (GPM) which perturbs the deuterium density and source profiles. By using absolutely calibrated experimental edge ionization profile measurements, radial profiles of diffusion ( D ) and convection ( v ) are calculated into the pedestal region without depending on modeling the edge ionization source. An analytic approach with closed-form expressions for the D and v profiles and a more advanced Bayesian approach show evidence of an inward particle convection on the order of 1 m s −1 extending to normalized poloidal flux ( Ψ N ) of 0.98. Meanwhile, diffusion reaches a minimum value of ( 0.03 ± 0.02 ) m 2 s −1 in the pedestal region. Notably, the Bayesian approach, which utilizes the Aurora 1.5 D forward model inside the IMPRAD OMFIT module, provides radially resolved transport profiles with associated uncertainty without requiring an explicit form for the perturbation to the density profile or source. The combination of experimental ionization measurements and Bayesian inference provides an enhanced robust framework for investigating edge particle transport coefficients to experimentally test transport physics in order to improve predictive capabilities in the tokamak edge.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

The chemical characterization of halo substructure in the Milky Way based on APOGEE

Galactic haloes in a Λ-CDM universe are predicted to host today a swarm of debris resulting from cannibalized dwarf galaxies. The chemodynamical information recorded in their stellar populations helps elucidate their nature, constraining the assembly history of the Galaxy. Using data from APOGEE and Gaia , we examine the chemical properties of various halo substructures, considering elements that sample various nucleosynthetic pathways. The systems studied are Heracles, Gaia -Enceladus/Sausage (GES), the Helmi stream, Sequoia, Thamnos, Aleph, LMS-1, Arjuna, I’itoi, Nyx, Icarus, and Pontus. Abundance patterns of all substructures are cross-compared in a statistically robust fashion. Our main findings include: (i) the chemical properties of most substructures studied match qualitatively those of dwarf Milky Way satellites, such as the Sagittarius dSph. Exceptions are Nyx and Aleph, which are chemically similar to disc stars, implying that these substructures were likely formed in situ ; (ii) Heracles differs chemically from in situ populations such as Aurora and its inner halo counterparts in a statistically significant way. The differences suggest that the star formation rate was lower in Heracles than in the early Milky Way; (iii) the chemistry of Arjuna, LMS-1, and I’itoi is indistinguishable from that of GES, suggesting a possible common origin; (iv) all three Sequoia samples studied are qualitatively similar. However, only two of those samples present chemistry that is consistent with GES in a statistically significant fashion; (v) the abundance patterns of the Helmi stream and Thamnos are different from all other halo substructures.

79 ASTRONOMY AND ASTROPHYSICS↗

Search for Dark Matter Ionization on the Night Side of Jupiter with Cassini

We present a new search for dark matter (DM) using planetary atmospheres. We point out that annihilating DM in planets can produce ionizing radiation, which can lead to excess production of ionospheric H 3 + . We apply this search strategy to the night side of Jupiter near the equator. The night side has zero solar irradiation, and low latitudes are sufficiently far from ionizing auroras, leading to a low-background search. We use data on ionospheric H 3 + emission collected three hours either side of Jovian midnight, during its flyby in 2000, and set novel constraints on the DM-nucleon scattering cross section down to about 10 − 38 cm 2 . We also highlight that DM atmospheric ionization may be detected in Jovian exoplanets using future high-precision measurements of planetary spectra. Published by the American Physical Society 2024

79 ASTRONOMY AND ASTROPHYSICS↗

Performance Evaluation of Heterogeneous GPU Programming Frameworks for Hemodynamic Simulations

Preparing for the deployment of large scientific and engineering codes on upcoming exascale systems with GPU-dense nodes is made challenging by the unprecedented diversity of device architectures and heterogeneous programming models. In this work, we evaluate the process of porting a massively parallel, fluid dynamics code written in CUDA to SYCL, HIP, and Kokkos with a range of backends, using a combination of automated tools and manual tuning. We use a proxy application along with a custom performance model to inform the results and identify additional optimization strategies. At scale performance of the programming model implementations are evaluated on pre-production GPU node architectures for Frontier and Aurora, as well as on current NVIDIA device-based systems Summit and Polaris. Real-world workloads representing 3D blood flow calculations in complex vasculature are assessed. Our analysis highlights critical trade-offs between code performance, portability, and development time.

Martin, Aristotle↗

In-Transit Data Transport Strategies for Coupled AI-Simulation Workflow Patterns

Coupled AI-Simulation workflows are becoming the major workloads for HPC facilities, and their increasing complexity necessitates new tools for performance analysis and prototyping of new in-situ workflows. We present SimAI-Bench, a tool designed to both prototype and evaluate these coupled workflows. In this paper, we use SimAI-Bench to benchmark the data transport performance of two common patterns on the Aurora supercomputer: a one-to-one workflow with co-located simulation and AI training instances, and a many-to-one workflow where a single AI model is trained from an ensemble of simulations. For the one-to-one pattern, our analysis shows that node-local and DragonHPC data staging strategies provide excellent performance compared Redis and Lustre file system. For the many-to-one pattern, we find that data transport becomes a dominant bottleneck as the ensemble size grows. Our evaluation reveals that file system is the optimal solution among the tested strategies for the many-to-one pattern.

Tummalapalli, Harikrishna [Argonne National Labora↗

HydraGNN v4.0

The new version of HydraGNN v4.0 provides additional core capabilities, such as: Inclusion of multi-body atomistic cluster expansion MACE, polarizable atom interaction neural network PAINN, and equivariant principal neighborhood aggregation (PNAEq) among the message passing layers supported -Inclusion of graph transformers to directly model long-range interactions between nodes that are distant in the graph topology Integration of graph transformers with message passing layers by combining the graph embedding generated by the two mechanisms, which allows for an improved expressivity of the HydraGNN architecture Improved re-implementation of multi-task learning (MTL) to allow its use for stabilized training across imbalanced, multi-source, multi-fidelity data Introduction of multi-task parallelism, a newly proposed type of model parallelism specifically for MTL architectures, which allows to dispatch different output decoding heads to different GPU devices Integration of multi-task parallelism with pre-existing distributed data parallelism to enable a 2D parallelization for distributed training Improved portability of the distributed training across Intel GPUs, which has been testes on ALCF exascale supercomputer Aurora Inclusion of 2-level fine-grained energy profilers portable across NVIDIA, AMD, and Intel GPUs to monitor the power and energy consumption associated with different functions executed by the HydraGNN code during data pre-load and training Restructuring of previous examples and inclusion of new sets of examples to illustrate the download, preprocess, and training of HydraGNN models on new large-scale open-source datasets for atomistic materials modeling (e.g., Alexandria, Transition1x, OMat24, OMol25)

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

HydraGNN v5.0

HydraGNN v5.0 expands the code base into a more portable, scalable, and flexible framework for scientific graph learning, with particular strength in atomistic machine-learning interatomic potentials and large-scale distributed training. The release adds Fully Sharded Data Parallel (FSDP) support alongside existing DDP and DeepSpeed paths, including FSDP-aware checkpointing and optimizer integration, and introduces a configurable multi-precision training workflow supporting FP32, BF16, and FP64 across GPUs and Intel XPUs. For atomistic modeling, HydraGNN v5.0 strengthens its MLIP capabilities through dynamic graph construction at every forward pass, energy-conserving force prediction via automatic differentiation, and per-atom energy loss formulations, while extending EGNN models to properly handle periodic boundary conditions. The release also broadens model expressiveness through graph-level attribute conditioning, adds new multi-task and model-parallel extensions such as MACE support and encoder/decoder branch optimization, and expands application coverage with integrated examples for datasets including OC25, Nabla2-DFT, QCML, Open Polymers 2026, and OPF. In parallel, HydraGNN v5.0 improves production readiness through performance optimizations for large-scale runs, stratified sampling and linear-regression preprocessing utilities, and tested installation scripts for DOE supercomputers including Frontier, Aurora, Perlmutter, and Andes. Overall, the release advances HydraGNN as a robust software platform for scalable graph neural networks across materials science, chemistry, and scientific machine learning workflows

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

Reducing Southern Ocean Shortwave Radiation Errors in the ERA5 Reanalysis with Machine Learning and 25 Years of Surface Observations

Earth system models struggle to simulate clouds and their radiative effects over the Southern Ocean, partly due to a lack of measurements and targeted cloud microphysics knowledge. We have evaluated biases of downwelling shortwave radiation in the ERA5 climate reanalysis using 25 years (1995–2019) of summertime surface measurements, collected on the Research and Supply Vessel (RSV) Aurora Australis, the Research Vessel (R/V) Investigator, and at Macquarie Island. During October–March daylight hours, the ERA5 simulation of SW down exhibited large errors (mean bias = 54 W m -2 , mean absolute error = 82 W m -2 , root-mean-square error = 132 W m -2 , and R 2 = 0.71). To determine whether we could improve these statistics, we bypassed ERA5’s radiative transfer model for SW down with machine learning–based models using a number of ERA5’s gridscale meteorological variables as predictors. These models were trained and tested with the surface measurements of SW down using a 10-fold shuffle split. An extreme gradient boosting (XGBoost) and a random forest–based model setup had the best performance relative to ERA5, both with a near complete reduction of the mean bias error, a decrease in the mean absolute error and root-mean-square error by 25% ± 3%, and an increase in the R 2 value of 5% ± 1% over the 10 splits. Large improvements occurred at higher latitudes and cyclone cold sectors, where ERA5 performed most poorly. We further interpret our methods using Shapley additive explanations. Our results indicate that data-driven techniques could have an important role in simulating surface radiation fluxes and in improving reanalysis products.

54 ENVIRONMENTAL SCIENCES↗

A graphics processing unit accelerated sparse direct solver and preconditioner with block low rank compression

We present the GPU implementation efforts and challenges of the sparse solver package STRUMPACK. The code is made publicly available on github with a permissive BSD license. STRUMPACK implements an approximate multifrontal solver, a sparse LU factorization which makes use of compression methods to accelerate time to solution and reduce memory usage. Multiple compression schemes based on rank-structured and hierarchical matrix approximations are supported, including hierarchically semi-separable, hierarchically off-diagonal butterfly, and block low rank. Here, in this paper, we present the GPU implementation of the block low rank (BLR) compression method within a multifrontal solver. Our GPU implementation relies on highly optimized vendor libraries such as cuBLAS and cuSOLVER for NVIDIA GPUs, rocBLAS and rocSOLVER for AMD GPUs and the Intel oneAPI Math Kernel Library (oneMKL) for Intel GPUs. Additionally, we rely on external open source libraries such as SLATE (Software for Linear Algebra Targeting Exascale), MAGMA (Matrix Algebra on GPU and Multi-core Architectures), and KBLAS (KAUST BLAS). SLATE is used as a GPU-capable ScaLAPACK replacement. From MAGMA we use variable sized batched dense linear algebra operations such as GEMM, TRSM and LU with partial pivoting. KBLAS provides efficient (batched) low rank matrix compression for NVIDIA GPUs using an adaptive randomized sampling scheme. The resulting sparse solver and preconditioner runs on NVIDIA, AMD and Intel GPUs. Interfaces are available from PETSc, Trilinos and MFEM, or the solver can be used directly in user code. We report results for a range of benchmark applications, using the Perlmutter system from NERSC, Frontier from ORNL, and Aurora from ALCF. For a high frequency wave equation on a regular mesh, using 32 Perlmutter compute nodes, the factorization phase of the exact GPU solver is about 6.5× faster compared to the CPU-only solver. The BLR-enabled GPU solver is about 13.8× faster than the CPU exact solver. For a collection of SuiteSparse matrices, the STRUMPACK exact factorization on a single GPU is on average 1.9× faster than NVIDIA’s cuDSS solver.

97 MATHEMATICS AND COMPUTING↗