Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel time integration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Skipper-in-CMOS: Nondestructive Readout With Subelectron Noise Performance for Pixel Detectors

The Skipper-in-CMOS image sensor integrates the nondestructive readout capability of skipper charge coupled devices (Skipper-CCDs) with the high conversion gain of a pinned photodiode (PPD) in a CMOS imaging process while taking advantage of in-pixel signal processing. This allows both single photon counting as well as high frame rate readout through highly parallel processing. The first results obtained from a ${15} \times {15}~\mu $ m2 pixel cell of a Skipper-in-CMOS sensor fabricated in Tower Semiconductor’s commercial 180-nm CMOS image sensor process are presented. Measurements confirm the expected reduction of the readout noise with the number of samples down to deep subelectron noise of $0.15\text {e}^ - $ , demonstrating the charge transfer operation from the PPD and the single photon counting operation when the sensor is exposed to light. This article also discusses new testing strategies employed for its operation and characterization.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Advanced architectures for high-performance quantum networking

As practical quantum networks prepare to serve an ever-expanding number of nodes, there has grown a need for advanced auxiliary classical systems that support the quantum protocols and maintain compatibility with the existing fiber-optic infrastructure. We propose and demonstrate a quantum local area network design that addresses current deployment limitations in timing and security in a scalable fashion using commercial off-the-shelf components. First, we employ White Rabbit switches to synchronize three remote nodes with ultra-low timing jitter, significantly increasing the fidelities of the distributed entangled states over previous work with Global Positioning System clocks. Second, using a parallel quantum key distribution channel, we secure the classical communications needed for instrument control and data management. Therefore, the conventional network that manages our entanglement network is secured using keys generated via an underlying quantum key distribution layer, preserving the integrity of the supporting systems and the relevant data in a future-proof fashion.

97 MATHEMATICS AND COMPUTING↗

Fast nonlinear iterative solver for an implicit, energy-conserving, asymptotic-preserving charged-particle orbit integrator

Here recently, an asymptotic-preserving (AP) particle orbit integrator has been proposed with remarkable properties including exact energy conservation, the ability to capture of all first-order drifts (including the ∇B-drift), the ability to capture trapped-passing boundaries with parallel velocity extremely close to the critical velocity, and the ability to transition from strongly to weakly magnetized spatial regions. The new AP orbit integrator is implicit, employing a Crank-Nicolson (CN) temporal discretization to ensure exact energy conservation. This, in turn, requires a local nonlinear iteration involving particle velocities and positions, and the local electromagnetic fields, to obtain the new-time solution. Ref. [1] did not attempt to provide an efficient solver for this system, and employed a brute-force GMRES-driven Jacobian-free Newton-Krylov (JFNK) solver to invert the particle orbit equations at every timestep for expediency. While JFNK is robust and reliable, it is also expensive and very intrusive for practical implementations of the method (it requires having the JFNK machinery available and solving a 6 x 6 Jacobian system iteratively once per iteration per particle).

97 MATHEMATICS AND COMPUTING↗

Galaxy bispectrum in the spherical Fourier-Bessel basis

The bispectrum, the three-point correlation in Fourier space, is a crucial statistic for studying many effects targeted by the next-generation galaxy surveys, such as primordial non-Gaussianity (PNG) and general relativistic (GR) effects on large scales. In this work we develop a formalism for the bispectrum in the spherical Fourier-Bessel (SFB) basis—a natural basis for computing correlation functions on the curved sky, as it diagonalizes the Laplacian operator in spherical coordinates. Working in the SFB basis allows for line-of-sight effects such as redshift space distortions and GR to be accounted for exactly, i.e., without having to resort to perturbative expansions to go beyond the plane-parallel approximation. Only analytic results for the SFB bispectrum exist in the literature given the intensive computations needed. We numerically calculate the SFB bispectrum for the first time, enabled by a few techniques: We implement a template decomposition of the redshift-space kernel Z 2 into Legendre polynomials, and separately treat the PNG and velocity-divergence terms. We derive an identity to integrate a product of three spherical harmonics connected by a Dirac delta function as a simple sum and use it to investigate the limit of a homogeneous and isotropic Universe. Furthermore, we present a formalism for convolving the signal with separable window functions and use a toy spherically symmetric window to demonstrate the computation and give insights into the properties of the observed bispectrum signal. While our implementation remains computationally challenging, it is a step toward a feasible full extraction of information on large scales via a SFB bispectrum analysis.

79 ASTRONOMY AND ASTROPHYSICS↗

Prediction of DIII-D Pedestal Structure From Externally Controllable Parameters

The sharp increase of pressure at the edge of a high confinement mode (H-mode) plasma, the pedestal, strongly impacts overall plasma performance. Predicting the pedestal is a necessity to control and optimize tokamak operations. Here, an experimental data-driven machine learning (ML) approach is presented that predicts the pedestal heights and widths of electron density (n e ) and electron temperature (T e ) profiles as well as the separatrix ne from externally controllable parameters such as the plasma shape, heating method and power, and gas puff rate and integrated gas puff. The OMFIT framework was used with DIII-D data to efficiently, robustly, and automatically build a database of pedestal parameters to train machine learning models. Database creation was enabled by the search engine tool for DIII-D data, TokSearch, which parallelizes data fetching, enabling fast searches through basic signals of thousands of DIII-D shots and selection of relevant time intervals. Principal Component Analysis (PCA) separated the database into three clusters that represent classes of plasma shapes that are regularly used in DIII-D. The most important parameters for setting the pedestal structure were plasma current (I p ), toroidal magnetic field (B Φ ), neutral beam heating power (P NBI ) and shaping quantities. The Deep Jointly Informed Neural Networks (DJINN) algorithm was applied to identify suitable neural network (NN) architectures that appropriately capture the features of the pedestal database. Separate NNs were implemented for each pedestal parameter, and ensembling methods were used to improve the prediction accuracy and allowed estimation of the prediction uncertainty. The pedestal predictions of the test dataset lie within the measurement uncertainties of the pedestal parameters. The NN outperformed simple Linear Regression (LR) analysis, indicating non-linear dependencies in the pedestal structure. The presented achievements illustrate a promising path for future research, using feature extraction to infer experimental trends and thereby improve pedestal models as well as deploying NN for a fast pedestal prediction in DIII-D scenario development.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Neural network methods for radiation detectors and imaging

Recent advances in image data proccesing through deep learning allow for new optimization and performance-enhancement schemes for radiation detectors and imaging hardware. This enables radiation experiments, which includes photon sciences in synchrotron and X-ray free electron lasers as a subclass, through data-endowed artificial intelligence. We give an overview of data generation at photon sources, deep learning-based methods for image processing tasks, and hardware solutions for deep learning acceleration. Most existing deep learning approaches are trained offline, typically using large amounts of computational resources. However, once trained, DNNs can achieve fast inference speeds and can be deployed to edge devices. A new trend is edge computing with less energy consumption (hundreds of watts or less) and real-time analysis potential. While popularly used for edge computing, electronic-based hardware accelerators ranging from general purpose processors such as central processing units (CPUs) to application-specific integrated circuits (ASICs) are constantly reaching performance limits in latency, energy consumption, and other physical constraints. These limits give rise to next-generation analog neuromorhpic hardware platforms, such as optical neural networks (ONNs), for high parallel, low latency, and low energy computing to boost deep learning acceleration (LA-UR-23-32395).

edge computing↗

Prediction of DIII-D Pedestal Structure from Externally Controllable Parameters

The sharp increase of pressure at the edge of a high confinement mode (H-mode) plasma, the pedestal, strongly impacts overall plasma performance. Predicting the pedestal is a necessity to control and optimize tokamak operations. An experimental data-driven machine learning (ML) approach is presented that predicts the pedestal heights and widths of electron density (ne) and electron temperature (Te) profiles as well as the separatrix ne from externally controllable parameters such as the plasma shape, heating method and power, and gas puff rate and integrated gas puff. The OMFIT framework was used with DIII-D data to efficiently, robustly, and automatically build a database of pedestal parameters to train machine learning models. Database creation was enabled by the search engine tool for DIII-D data, TokSearch, which parallelizes data fetching, enabling fast searches through basic signals of thousands of DIII-D shots and selection of relevant time intervals. Principal Component Analysis (PCA) separated the database into three clusters that represent classes of plasma shapes that are regularly used in DIII-D. The most important parameters for setting the pedestal structure were plasma current (Ip), toroidal magnetic field (Bφ), neutral beam heating power (PNBI) and shaping quantities. The Deep Jointly Informed Neural Networks (DJINN) algorithm was applied to identify suitable neural network (NN) architectures that appropriately capture the features of the pedestal database. Separate NNs were implemented for each pedestal parameter, and ensembling methods were used to improve the prediction accuracy and allowed estimation of the prediction uncertainty. The pedestal predictions of the test dataset lie within the measurement uncertainties of the pedestal parameters. The NN outperformed simple Linear Regression (LR) analysis, indicating non-linear dependencies in the pedestal structure. The presented achievements illustrate a promising path for future research, using feature extraction to infer experimental trends and thereby improve pedestal models as well as deploying NN for a fast pedestal prediction in DIII-D scenario development.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Low-synch Gram–Schmidt with delayed reorthogonalization for Krylov solvers

The parallel strong-scaling of iterative methods is often determined by the number of global reductions at each iteration. Low-synch Gram-Schmidt algorithms are applied here to the Arnoldi algorithm to reduce the number of global reductions and therefore to improve the parallel strong-scaling of iterative solvers for nonsymmetric matrices such as the GMRES and the Krylov-Schur iterative methods. In the Arnoldi context, the factorization is "left-looking" and processes one column at a time. Among the methods for generating an orthogonal basis for the Arnoldi algorithm, the classical Gram-Schmidt algorithm, with reorthogonalization (CGS2) requires three global reductions per iteration. A new variant of CGS2 that requires only one reduction per iteration is presented and applied to the Arnoldi algorithm. Delayed CGS2 (DCGS2) employs the minimum number of global reductions per iteration (one) for a one-column at-a-time algorithm. The main idea behind the new algorithm is to group global reductions by rearranging the order of operations. DCGS2 must be carefully integrated into an Arnoldi expansion or a GMRES solver. Numerical stability experiments assess robustness for Krylov-Schur eigenvalue computations. Performance experiments on the ORNL Summit supercomputer then establish the superiority of DCGS2 over CGS2.

97 MATHEMATICS AND COMPUTING↗

Programming approaches for scalability, performance, and portability of combustion physics codes

Here, this paper presents the process, strategy, and results associated with porting a typical combustion physics flow solver to current state-of-the-art and future massively-parallel computer architectures. Major focus is placed on the distinct algorithmic structure of these types of codes and how it can be integrated with modern programming paradigms for heterogeneous platforms (i.e., distributed many-core systems with accelerators). An end-to-end case study is presented that exemplifies the process in a generic manner, which then serves as a clear guide with respect to the strategy and best practices leading to a robust and adaptable framework that performs well, is durable over time, is portable, and requires minimal human-effort. This end is accomplished beginning with the use of a mature, validated, structured, multiblock code framework optimized for application of both Large Eddy Simulation (LES) and Direct Numerical Simulation (DNS). This code has been ported to a variety of platforms over the past decade, including most recently the Oak Ridge Leadership Computing Facility’s “Summit” Platform. The experience gained on these multiple platforms provides general insights and thus the results presented are not specific to any one code or platform other than the overarching trend toward distributed many-core systems with accelerators in order to move toward exascale performance. The resultant performance and scalability of the ported code is demonstrated on a real-world application; a state-of-the-art rotating detonation rocket engine simulation that matches the complex geometry and boundary conditions imposed as part of a companion experimental campaign.

97 MATHEMATICS AND COMPUTING↗

ParFlow Sand Tank: A tool for groundwater exploration

The ParFlow Sand Tank model is an open source application designed to allow users to interactively simulate and visualize groundwater movement through the subsurface. The app is designed for both research and education; teaching hydrogeology concepts and making it easy explore and run sophisticated groundwater simulations. Our goal is to support increased accessibility and usability of research grade hydrology tools for research and teaching. The Sand Tank application simulates groundwater and surface water fluxes as well as contaminant transport in real time using the integrated physical hydrology model ParFlow (Kollet & Maxwell, 2006; Maxwell & Miller, 2005; Osei-Kuffuor et al., 2014) and the particle tracking code EcoSlim (Maxwell et al., 2019). ParFlow is a numerical hydrology model that simulates spatially distributed groundwater and surface water flow. It is a well established research tool with more than 90 publications documenting its development use to advance our understanding of groundwater dynamics and groundwater surface water interactions from the hillslope to the continental scale e.g. (Condon et al., 2020; Condon & Maxwell, 2019; Maxwell & Condon, 2016). It is designed for efficient parallel computation and has been run on many platforms spanning from laptops to supercomputers. However, one of the challenges of ParFlow is that it requires significant training and hydrologic expertise to develop simulations. The Sand Tank application makes this model accessible to anyone for education and exploration. Our application uses ParFlow for its simulation backend and ParaView for the data loading and processing. The communication infrastructure relies on the ParaViewWeb framework. We use model templates deployed in Docker images to setup the Sand Tank framework. Users can build the application locally or interact with it through our web deployment. When interacting with a template users can interactively change model parameters like subsurface processes or pump/inject water into the subsurface and watch the system respond to their changes in real time as the simulation runs. Additionally, our template setup will allow more advanced users to build custom templates of increasing complexity for both research and educational purposes.

54 ENVIRONMENTAL SCIENCES↗

Active Learning for Metamaterial Optimization on HPC and QC Integrated Systems

Active learning algorithms, integrating machine learning, quantum computing and optics simulation in an iterative loop, offer a promising approach to optimizing metamaterials. However, these algorithms can face difficulties in optimizing highly complex structures due to computational limitations. High-performance computing (HPC) and quantum computing (QC) integrated systems can address these issues by enabling parallel computing. In this study, we develop an active learning algorithm working on HPC-QC integrated systems. We evaluate the performance of optimization processes within active learning (i.e., training a machine learning model, problem-solving with quantum computing, and evaluating optical properties through wave-optics simulation) for highly complex metamaterial cases. Our results showcase that utilizing multiple cores on the integrated system can significantly reduce computational time, thereby enhancing the efficiency of optimization processes. Therefore, we expect that leveraging HPC-QC integrated systems helps effectively tackle large-scale optimization challenges in general.

Kim, Seongmin↗

BISON Robustness and Performance Improvements

BISON is a modern finite-element based nuclear fuel performance code that has been under development at the Idaho National Laboratory (USA) since 2009 [1]. The code is applicable to both steady and transient fuel behavior and can be used to analyze 1D (spherically symmetric), 2D (axisymmetric and generalized plane strain) or 3D geometries. BISON is the fuel performance code used within CASL for LWR fuel under both normal operating and accident conditions. BISON is built using the INL Multiphysics ObjectOriented Simulation Environment, or MOOSE [2, 3]. MOOSE is a massively parallel, finite element-based framework to solve systems of coupled non-linear partial differential equations using the Jacobian-Free Newton Krylov (JFNK) method [4]. This enables investigation of computationally large problems, for example a full stack of discrete pellets in a LWR fuel rod, or every rod in a full reactor core. MOOSE supports the use of complex two and three-dimensional meshes and uses implicit time integration, important for the widely varied time scale in nuclear fuel simulation. An object-oriented architecture is employed which greatly minimizes the programming effort required to add new material and behavioral models. The flexibility of the implicit and fully coupled multiphysics approach comes with a need for constructing suitable approximations for the Jacobian matrix of the coupled system used for either preconditioning a Krylov solve or in a direct Newton solve. Preconditioning options for Bison problems need to be revisited with new preconditioning methods becoming available.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Nonlinear simulations of GAEs in NSTX-U

A set of nonlinear simulations has been performed in order to study the nonlinear evolution of unstable global Alfvén eigenmodes in the National Spherical Torus Experiment-Upgrade (NSTX-U). Results of the single toroidal mode number, n, simulations are compared with a full nonlinear simulation (all toroidal harmonics included). In single-n simulations, the conservation of two integrals of motion of a particle in a cyclotron resonance with a monochromatic wave is demonstrated, resulting in a one-dimensional evolution of the particle distribution in (E,μ,pϕ) phase-space. Nonlinear simulations (both single-n and full nonlinear) show a significant redistribution of the resonant fast ions, especially in the pitch parameter. Thus, the changes in the resonant particle's parallel and perpendicular energies can be several times larger than the total particle energy change, with only a small fraction transferring into the excitation of the mode itself. This implies that even a relatively small amplitude mode can significantly modify the beam distribution in the resonant region. For the NSTX-U case considered, the single-n simulation results are close to full nonlinear simulation only for the most unstable mode, in which case the saturation amplitudes and changes in the fast ion distribution are comparable. In contrast, peak amplitudes of subdominant modes in all-n simulations are smaller by a factor of 3–10 compared to single-n runs due to the flattening of the beam ion distribution by the fastest growing mode.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Summer Internship Report: ARA2 Benchmarking

Over the past decade, the RISC-V Instruction Set Architecture (ISA) has emerged as a significant player in both academic and industrial processor design due to its open-source nature, modular extension system, and versatility across domains ranging from microcontrollers to high-performance computing (HPC). One of its most important recent advancements is the RISC-V Vector Extension (RVV), which enables explicit data-level parallelism through vector registers and vectorized instructions. Unlike traditional SIMD (Single Instruction, Multiple Data) architectures that fix vector lengths at design time, RVV uses the concept of VLEN (vector register length) as a hardware-independent parameter and allows software to adapt dynamically to the available vector width. This flexible approach ensures portability across implementations while enabling scalable performance. The ARA2 core is a parameterizable RISC-V vector processor developed at the Integrated Systems Lab at ETH Zürich and the University of Bologna. Designed as a tightly-coupled accelerator to a scalar RISC-V core, ARA2 implements the RVV 1.0 specification and offers tunable architectural parameters such as the number of vector lanes, VLEN, and cache sizes.

97 MATHEMATICS AND COMPUTING↗

ICED: An Integrated CGRA Framework Enabling DFVS-Aware Acceleration

oarse-grained reconfigurable arrays (CGRAs) are a promising solution to enable energy-efficient acceleration of applications from different domains. By leveraging reconfiguration at the functional level, they can adapt to significantly different computational patterns. Existing CGRA mapping approaches extract instruction-level parallelism, exploit loop-pipelining opportunities, guarantee the data dependency, and target high throughput of a given loop. However, the recurrence data-dependency in the DFG and the mismatch between required and available computing/communication resources complicate the mapping, and might lead to significant unbalances in the utilization of the CGRA's tiles. This results in wasted power for tiles with low utilization. Applying dynamic voltage and frequency scaling (DVFS) can potentially solve this challenge and improve energy efficiency by adjusting voltage and frequency of different tiles independently. CGRAs have also been successful in accelerating data-dependent streaming applications. However, in these applications, the execution time of each kernel in the pipeline might dynamically vary depending on the characteristics of the input. This also leads to under-utilization of resources for the dynamically changing kernels that do not limit the application throughput. DVFS can also improve energy efficiency for these applications by dynamically changing the voltage and frequency levels of tiles that host non performance-constraining kernels. This paper proposes ICEDTEA -- an integrated DVFS-aware framework to map applications on CGRAs that support power islands. ICEDTEA proposes a CGRA architecture supporting DVFS islands at varying granularity (from a single tile to a group of tiles) and the related DVFS-aware compilation and mapping toolchain. ICEDTEA is the first work that introduces DVFS support for spatio-temporal CGRAs at power-island levels. The experimental evaluation shows that ICEDTEA improves average utilization by 2.3$\times$ and energy-efficiency by 1.32$\times$ over a conventional CGRA. With streaming applications, ICEDTEA improves energy efficiency by 1.12$\times$ over a state-of-the-art CGRA that introduces partial dynamic reconfiguration to adapt to variations in kernels' throughput.

Tan, Cheng↗

Dissolved oxygen sensor in an automated hyporheic sampling system reveals biogeochemical dynamics

Many river corridor systems frequently experience rapid variations in river stage height, hydraulic head gradients, and residence times. The integrated hydrology and biogeochemistry of such systems is challenging to study, particularly in their associated hyporheic zones. Here we present an automated system to facilitate 4-dimensional study of dynamic hyporheic zones. It is based on combining real-time in-situ and ex-situ measurements from sensor/sampling locations distributed in 3-dimensions. A novel dissolved oxygen (DO) sensor was integrated into the system during a small scale study. We measured several biogeochemical and hydrologic parameters at three subsurface depths in the riverbed of the Columbia River in Washington State, USA, a dynamic hydropeaked river corridor system. During the study, episodes of significant DO variations (~+/- 4 mg/l) were observed, with minor variation in other parameters (e.g., <~+/-0.15 mg/l NO 3 ). DO concentrations were related to hydraulic head gradients, showing both hysteretic and non-hysteretic relationships with abrupt (hours) transitions between the two types of relationships. The observed relationships provide a number of hypotheses related to the integrated hydrology and biogeochemistry of dynamic hyporheic zones. We suggest that preliminary high-frequency monitoring is advantageous in guiding the design of long term monitoring campaigns. The study also demonstrated the importance of measuring multiple parameters in parallel, where the DO sensor provided the key signal for identifying/detecting transient phenomena.

54 ENVIRONMENTAL SCIENCES↗

Bioresorbable Primary Battery Anodes Built on Core–Double-Shell Zinc Microparticle Networks

Bioresorbable implantable electronics require power sources that are also bioresorbable with controllable electrical output and lifetime. In this paper, we report a bioresorbable zinc primary battery anode filament based on a zinc microparticle (MP) network coated with chitosan and Al 2 O 3 double shells. When discharged in 0.9% NaCl saline, a Zn MP filament with a 0.17 × 2 mm 2 cross-sectional area exhibited a stable voltage output of 0.55 V at a current of 0.01 mA. Covered by chitosan and Al 2 O 3 double shells, the zinc MP filament exhibited a directional dissolution behavior with a tunable lifetime approximately linear to its length. A stable 200 h discharging time was achieved with a 15 mm Zn MP filament. The maximum output power was found to be 12 μW at 0.03 mA for one filament. The linearity relationship between the current output and the filament cross-sectional area suggested a facile strategy to raise the power output at constant discharging voltage. The filaments could also be connected in series and in parallel to boost its overall voltage and current output, demonstrating their excellent integration capability. Furthermore, this work presents a promising pathway toward bioresorbable transient batteries with controllable lifetime and power output, demonstrating a great potential for powering transient implantable biomedical devices.

25 ENERGY STORAGE↗

Snow ALbedo eVOlution (SALVO) Campaign Spectral Albedo and Related Measurements from April - June, 2024 in Utqiagivk, AK

A field-portable spectroradiometer, referred to herein as an ‘ASD’, was used to make spatially-distributed spectral albedo (350 – 2500 nm) measurements on tundra and sea ice surfaces. The ASD detector is carried in a backpack and controlled via a computer mounted on the front of the operator (see Figure 1). The ASD measures the spectral irradiance from a fiber optic cable that is routed from the backpack to a custom, gooseneck cosine collector mounted on the end of a 1-m long boom (Grenfell and Perovich, 2008). The boom was held at hip height (approximately 1 m) and had an integrated bubble level for levelling. To make an albedo measurement, first the operator collect an incident (down-welling) irradiance, followed by a reflected (up-welling) measurement. The time between incident and reflected measurements was typically between 11 and 26 seconds (interquartile range). For each measurement, 10 spectra are averaged together. Albedo is calculated as the ratio of the reflected to incident measurement, which obviates the need for absolute radiometric calibration. Albedo measurements were taken parallel to the 200-m albedo lines at 5-m increments (41 measurements) ~1 m south of the line. While the ASD operator was making measurements, an assistant kept notes on the scan number associated with each measurement, the surface type (see below), and collected photos of each measurement (see companion oblique photos data archive). Measurements were made within 3 hours of solar noon.

ASD Spectroradiometer↗