Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Making ultrastrong steel tough by grain-boundary delamination

Developing ultrahigh-strength steels that are ductile, fracture resistant, and cost effective would be attractive for a variety of structural applications. We show that improved fracture resistance in a steel with an ultrahigh yield strength of nearly 2 gigapascals can be achieved by activating delamination toughening coupled with transformation-induced plasticity. Delamination toughening associated with intensive but controlled cracking at manganese-enriched prior-austenite grain boundaries normal to the primary fracture surface dramatically improves the overall fracture resistance. As a result, fracture under plane-strain conditions is automatically transformed into a series of fracture processes in “parallel” plane-stress conditions through the thickness. The present “high-strength induced multidelamination” strategy offers a different pathway to develop engineering materials with ultrahigh strength and superior toughness at economical materials cost.

36 MATERIALS SCIENCE↗

Using Parameter Sweep in WaterTAP to Analyze New Water Treatment Technologies

We describe a powerful and generalized parameter sweep tool in this report that was originally developed to analyze the performance of existing and novel water treatment models being developed in WaterTAP. Since WaterTAP is built upon IDAES and Pyomo, the parameter sweep tool can be used to systematically explore and debug the behavior of most Pyomo and IDAES numerical models. In order to enable meaningful analyses, the parameter sweep tool has been designed with the following features: 1) Model flexibility: The parameter sweep tool does not enforce any restrictions on the types of models that can be used with it. As long as a Pyomo model can be solved and the parameter is active and mutable, the tool only needs functions that describe how to run the model, the sweep parameters, and the output quantities of interest. 2) Flexible sampling: The parameter sweep tool has inbuilt functions to generate samples from a random distribution or a multidimensional Euclidean space. Furthermore, the users have to ability to supply samples generated from a tool of their choice. 3) Multiple sweep types: A user can choose from one of 3 types of parameter sweeps depending on their needs. 4) Detailed outputs: Outputs generated by the parameter sweep tool can be stored in detailed H5 file or user-friendly CSV files for post processing. 5) Parallel computing: The parameter sweep supports shared and distributed memory parallel computing to enable the use of high performance computers (HPC) for large-scale analyses. 6) Modular: The parameter sweep tool is self-contained and can easily be integrated within an outer-loop analysis or as desired by the user. 7) Ease of use: The tool is well documented and a simple sweep can be easily executed by following the online documentation in a few lines of code. We demonstrate the use of the parameter sweep tool on a simple water treatment system from the WaterTAP repository and show its parallel scaling performance on an Apple laptop and NREL's Eagle HPC. The parameter sweep tool is actively being used with models currently being developed within WaterTAP and we expect its use to grow beyond it to other IDAES and Pyomo models.

97 MATHEMATICS AND COMPUTING↗

Graph-Learning-Assisted State and Event Tracking for Solar-Penetrated Power Grids with Heterogeneous Data Sources

Unlike transmission systems, distribution systems do not typically contain sufficient metering to enable real-time state estimation. The lack of sufficient real-time measurements prohibits accurate and timely monitoring of the state of distribution systems. As a result, control and optimal operation of distribution systems, especially those containing large numbers of renewable generation units are not possible without proper data and information about the current state of the system. The main motivation of this project is to address this shortcoming by developing an approach which provides “predicted” real-time measurements so that they can be used to execute a distribution system state estimator. Thus, the objective of the project is to make the distribution systems fully observable, such that the hosting capacity for solar generation can be accurately estimated, and unnecessary solar curtailments can be avoided. In order to accomplish this goal, the project investigated the use of a grid-model-informed machine learning (ML) tool which integrates heterogeneous data streams obtained from AMI meters, SCADA as well as PMU measurements and created synchronous measurement snapshots for the state estimator (SE); and developed a hybrid robust SE which provides not only accurate state estimates but also real-time feedback for the ML model refinement.

14 SOLAR ENERGY↗

Massively parallel modeling and inversion of electrical resistivity tomography data using PFLOTRAN

Abstract. Electrical resistivity tomography (ERT) is a broadly accepted geophysical method for subsurface investigations. Interpretation of field ERT data usually requires the application of computationally intensive forward modeling and inversion algorithms. For large-scale ERT data, the efficiency of these algorithms depends on the robustness, accuracy, and scalability on high-performance computing resources. In this regard, we present a robust and highly scalable implementation of forward modeling and inversion algorithms for ERT data. The implementation is publicly available and developed within the framework of PFLOTRAN, an open-source, state-of-the-art massively parallel subsurface flow and transport simulation code. The forward modeling is based on a finite-volume discretization of the governing differential equations, and the inversion uses a Gauss–Newton optimization scheme. To evaluate the accuracy of the forward modeling, two examples are first presented by considering layered (1D) and 3D earth conductivity models. The computed numerical results show good agreement with the analytical solutions for the layered earth model and results from a well-established code for the 3D model. Inversion of ERT data, simulated for a 3D model, is then performed to demonstrate the inversion capability by recovering the conductivity of the model. To demonstrate the parallel performance of PFLOTRAN's ERT process model and inversion capabilities, large-scale scalability tests are performed by using up to 131 072 processes on a leadership class supercomputer. These tests are performed for the two most computationally intensive steps of the ERT inversion: forward modeling and Jacobian computation. For the forward modeling, we consider models with up to 122 ×106 degrees of freedom (DOFs) in the resulting system of linear equations and demonstrate that the code exhibits almost linear scalability on up to 10 000 DOFs per process. On the other hand, the code shows superlinear scalability for the Jacobian computation, mainly because all computations are fairly evenly distributed over each process with no parallel communication.

58 GEOSCIENCES↗

Performance enhancement of planar MOC calculation in hexagonal geometries by assembly-wise domain decomposition

The parallelization of planar Method of Characteristics (MOC) calculation with the higher order scattering in the hexagonal version of the nTRACER direct whole core calculation code is enhanced by adopting assembly-wise domain decomposition and other techniques. In order to augment the hexagonal domain decomposition, the three-color scheme, the angular decomposition, and the angular flux storage scheme are examined. The parallel performance assessed for the VVER benchmark problems demonstrates that the hexagonal assembly-wise domain decomposition significantly reduces the computing time in ray tracing by up to 64 %. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Parallel transport sweeps on two-dimensional cartesian and hexagonal grids

This paper aims to provide a proof of concept for parallel transport sweeps on two-dimensional hexagonal grids for the discrete ordinates transport equation. While the method is an extension of the popular and well-established Koch-Baker-Alcoulffe (KBA) algorithm, there are significant differences between the cartesian and hexagonal grid and thereafter sweep. The most important is the three-way connectivity of hexagons within the grid which creates greater dependencies between the elements. The KBA method in structured orthogonal grids was first implemented in the DRAGON5 code and the method is first described here. The differences in implementation for the hexagonal grid are also described. Benchmark results are also presented, showing roughly 10 times speedup in computational times with roughly 100 processors, in both cases. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Extension of the high-resolution thermal-hydraulics code ESCOT to hexagonal core geometries for multi-physics calculations

The extension of the capabilities of the pin-level nuclear reactor core thermal-hydraulics (T/H) code ESCOT to analyze hexagonal fueled cores and its performance are presented. ESCOT is an accurate yet fast core thermal-hydraulics solution aiming at high-fidelity and high-resolution multi-physics core analysis in the framework of massively parallel computing platforms. Its algorithm solution is based on the four-equation drift-flux model for two-phase calculations, these are numerically solved by applying the Finite Volume Method (FVM) and the Semi-Implicit Method for Pressure-Linked Equation (SIMPLE)-like algorithm in a staggered grid system. Constitutive models such as turbulent mixing, pressure drop, and vapor generation are employed to simulate key phenomena in subchannel-scale analysis. ESCOT is parallelized by a double (radial and axial) domain decomposition that enables its highly parallelized execution. The coupling of the code with the neutronics whole core solver for hexagonal geometries nTRACER is described. The newly implemented ESCOT features are validated by comparing single assembly and full core steady state nTRACER-ESCOT solutions with nTRACER standalone internal one-dimensional T/H solver results. The validation problems are based on the VVER 440 and VVER 1000 cores. ESCOT results show differences within an acceptable range with respect to the simple 1D nTRACER built-in solver. (authors)

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Enhanced Sustainability in Treatment of Contaminated Metals for Clearance and Recycling - 20290

Cyclife Sweden AB has, for more than 30 years, provided metal treatment services to the domestic and international nuclear industry with the aim of creating value through metal clearance and recycling. These operations are to a large extent based on the concept that the materials, after pre-treatment, including decontamination, will be melted for clearance. During the melting process, some nuclides separate from the metal while the remaining radioactivity becomes fully homogenized in the metal bath and the resulting ingots, along with the samples taken before casting At a time when activities to fight Climate Challenge are in focus, any savings in energy consumption, in addition to the positive environmental effect of recycling metals back to industry, is more than welcome. If the clearance of certain metallic items can be achieved safely, without using melting, it is the environmentally preferred option provided it carries a reasonable cost. It is even better if an object can go through a decontamination and clearance process prior to re-use inside or outside the nuclear sector. Cyclife Sweden is currently investing in a direct clearance process, to be operated in parallel with the melting process. This new process is expected to be fully operational in the late summer of 2020 and includes pre-treatment, decontamination and heavily shielded clearance measurement installations. The treatment will also allow clearance measurements of larger objects without the need for segmentation, along with allowing objects composed of different materials, like electric motors can be accepted for treatment. This paper gives an overview of the circular economy aspects in a nuclear back-end context, the Cyclife approach to further implement a circular economy approach in the processes, Material Acceptance Criteria for treatment, the facility set-up for the new process and how the melting route and the direct route can be combined in an efficient industrial scale process. To put the new capabilities into perspective, the two different treatment and clearance approaches, i.e. direct clearance and melting, are to be compared. (authors)

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Metaplastic and energy-efficient biocompatible graphene artificial synaptic transistors for enhanced accuracy neuromorphic computing

CMOS-based computing systems that employ the von Neumann architecture are relatively limited when it comes to parallel data storage and processing. In contrast, the human brain is a living computational signal processing unit that operates with extreme parallelism and energy efficiency. Although numerous neuromorphic electronic devices have emerged in the last decade, most of them are rigid or contain materials that are toxic to biological systems. In this work, we report on biocompatible bilayer graphene-based artificial synaptic transistors (BLAST) capable of mimicking synaptic behavior. The BLAST devices leverage a dry ion-selective membrane, enabling long-term potentiation, with ~50 aJ/µm 2 switching energy efficiency, at least an order of magnitude lower than previous reports on two-dimensional material-based artificial synapses. The devices show unique metaplasticity, a useful feature for generalizable deep neural networks, and we demonstrate that metaplastic BLASTs outperform ideal linear synapses in classic image classification tasks. With switching energy well below the 1 fJ energy estimated per biological synapse, the proposed devices are powerful candidates for bio-interfaced online learning, bridging the gap between artificial and biological neural networks.

97 MATHEMATICS AND COMPUTING↗

Surrogate optimization of variational quantum circuits

Variational quantum eigensolvers are touted as a near-term algorithm capable of impacting many applications. However, the potential has not yet been realized, with few claims of quantum advantage and high resource estimates, especially due to the need for optimization in the presence of noise. Finding algorithms and methods to improve convergence is important to accelerate the capabilities of near-term hardware for VQE or more broad applications of hybrid methods in which optimization is required. To this goal, we look to use modern approaches developed in circuit simulations and stochastic classical optimization, which can be combined to form a surrogate optimization approach to quantum circuits. Using an approximate (classical CPU/GPU) state vector simulator as a surrogate model, we efficiently calculate an approximate Hessian, passed as an input for a quantum processing unit or exact circuit simulator. This method will lend itself well to parallelization across quantum processing units. We demonstrate the capabilities of such an approach with and without sampling noise and a proof-of-principle demonstration on a quantum processing unit utilizing 40 qubits.

Gustafson, Erik J. [RIACS, Mtn. View] (ORCID:00000↗

Proximity Portability and in Transit , M-to-N Data Partitioning and Movement in SENSEI [Book Chapter]

In high-performance parallel in situ processing, the term in transit processing refers to those configurations where data must move from a producer to a consumer that runs on separate resources. In the context of parallel and distributed computing on an HPC platform one of the central challenges is to determine a mapping of data from producer ranks to consumer ranks. This problem is complicated by the heterogeneity that arises in producer-consumer pairs, such as when producer and consumer codes have different levels of concurrency, different scaling characteristics, or different data models. The resulting mapping and movement of data from M producer to N consumer ranks can have a significant impact on aggregate application performance, particularly when the data consumer requires only a subset of the overall data for its task. This chapter focuses on the design considerations that underlie SENSEI’s implementation to this challenging problem. These design considerations extend the core SENSEI architecture and include ideas like the need to accommodate flexibility in the choice of different partitioning methods, the ability for a data consumer to request and receive only the subset of data needed for its particular operation, and the ability to leverage any of several different data transport tools. The idea of proximity portability, being able to use different data transport methods as part of an in transit workflow, is illustrated through the use of three different transport layers where switching from one transport tool to another is accomplished with only a configuration file change. Here, the chapter also includes a performance analysis summary showing the performance gains that are possible in terms of multiple metrics, such as memory footprint, time to solution, and amount of data moved, when using optimized partitioners in an in transit setting, gains that are made possible by the implementation shaped by specific design considerations.

Bethel, E. Wes↗

Space-Time Block Preconditioning for Incompressible Flow

Parallel-in-time methods have become increasingly popular in the simulation of time-dependent numerical PDEs, allowing for the efficient use of additional message passing interface processes when spatial parallelism saturates. Most methods treat the solution and parallelism in space and time separately. In contrast, all-at-once methods solve the full space-time system directly, largely treating time as simply another spatial dimension. All-at-once methods offer a number of benefits over separate treatment of space and time, most notably significantly increased parallelism and faster time to solution (when applicable). However, the development of fast, scalable all-at-once methods has largely been limited to time-dependent (advection-)diffusion problems. This paper introduces the concept of space-time block preconditioning for the all-at-once solution of incompressible flow. By extending well-known concepts of spatial block preconditioning to the space-time setting, we develop a block preconditioner whose application requires the solution of a space-time (advection-)diffusion equation in the velocity block, coupled with a pressure Schur complement approximation consisting of independent spatial solves at each time-step, and a space-time matrix-vector multiplication. The new method is tested on four classical models in incompressible flow. Finally, the results indicate perfect scalability in refinement of spatial and temporal mesh spacing, perfect scalability in nonlinear Picard iteration count when applied to a nonlinear Navier--Stokes problem, and minimal overhead in terms of number of preconditioner applications compared with sequential time-stepping.

97 MATHEMATICS AND COMPUTING↗

Identification and mitigation of memory block timing issue in ITk ABCStar during ASIC production

The ABCStar is a mixed-signal front-end readout ASIC for the strips sensor portion of the ATLAS ITk detector being developed as part of the High-Luminosity LHC upgrade. In pre-production testing, a subtle design flaw was uncovered in the ABCStar that was reducing wafer yields in some manufactured lots from the expected 90% to as low as 2%. The root cause was determined to be a timing issue in the logic synthesized to control previously silicon proven memory blocks re-used for this ASIC. The solutions proposed included manufacturing process changes by the wafer foundry, changes to the operating parameters for the ABCStar in the detector, and the possibility that a redesign might be required. The two mitigation efforts were undertaken in parallel, with the process modification route a less desirable solution since already manufactured wafers would need to be scrapped in favour of the new ones. Based on a knowledge of the existing process, and testing done on the worst performing wafers, it was proposed that raising the core operating voltage of the ABCStar from 1.20V to 1.25V could address the timing issue by sufficiently speeding up its transistors. An extensive testing program that included the effects of temperature and radiation expected over the lifetime of the ITk detector was conducted to validate that approach. Those tests and studies proved that even the worst performing wafers would have yields over 80% with the 1.25V core voltage, and neither the modified process nor redesign would be required for ensuring reliable operation of the ITk. Based on testing, a further timing mitigation was implemented to provide an additional margin of reliability by increasing the duty cycle of the clock to the ABCStar. Testing of all ABCStar wafers has been completed and the production of the detector modules using these ASICs is now well underway as a result of the efforts detailed herein.

FOS: Physical sciences↗

ORBIT-2: Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling

Sparse observations and coarse-resolution climate models limit effective regional decision-making, underscoring the need for robust downscaling. However, existing AI methods struggle with generalization across variables and geographies and are constrained by the quadratic complexity of Vision Transformer (ViT) self-attention. We introduce ORBIT-2, a scalable foundation model for global, hyper-resolution climate downscaling. ORBIT-2 incorporates two key innovations: (1) Residual Slim ViT (Reslim), a lightweight architecture with residual learning and Bayesian regularization for efficient, robust prediction; and (2) TILES, a tile-wise sequence scaling algorithm that reduces self-attention complexity from quadratic to linear, enabling long-sequence processing and massive parallelism. ORBIT-2 scales to 10 billion parameters across 65,536 GPUs, achieving up to 4.1 ExaFLOPS sustained throughput and 74–98% strong scaling efficiency. It supports downscaling to 0.9 km global resolution and processes sequences up to 4.2 billion tokens. On 7 km resolution benchmarks, ORBIT-2 achieves high accuracy with R2 scores in range of 0.98–0.99 against observation data.

Wang, Xiao [ORNL] (ORCID:0000000165451943)↗

ORBIT-2 Weather and Climate Downscaling Software Repository

ORBIT-2 is a scalable foundation model for global, hyper-resolution climate and weather downscaling. ORBIT-2 incorporates two key innovations: (1) Residual Slim ViT (Reslim), a lightweight architecture with residual learning and Bayesian regularization for efficient, robust prediction; and (2) TILES, a tile-wise sequence scaling algorithm that reduces self-attention complexity from quadratic to linear, enabling long-sequence processing and massive parallelism. ORBIT-2 scales to 10 billion parameters across 65,536 GPUs, achieving up to 4.1 ExaFLOPS sustained throughput and 74–98% strong scaling efficiency. It supports downscaling to 0.9 km global resolution and processes sequences up to 4.2 billion tokens. On 7 km resolution benchmarks, ORBIT-2 achieves high accuracy with 𝑅2 scores in range of 0.98–0.99 against observation data.

Wang, Xiao [Oak Ridge National Laboratory]↗