Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed and parallel processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Newly Released Capabilities in the Distributed-Memory SuperLU Sparse Direct Solver

We present the new features available in the recent release of SuperLU_DIST, Version 8.1.1. SuperLU_DIST is a distributed-memory parallel sparse direct solver. The new features include (1) a 3D communication-avoiding algorithm framework that trades off inter-process communication for selective memory duplication, (2) multi-GPU support for both NVIDIA GPUs and AMD GPUs, and (3) mixed-precision routines that perform single-precision LU factorization and double-precision iterative refinement. Apart from the algorithm improvements, we also modernized the software build system to use CMake and Spack package installation tools to simplify the installation procedure. Throughout the article, we describe in detail the pertinent performance-sensitive parameters associated with each new algorithmic feature, show how they are exposed to the users, and give general guidance of how to set these parameters. We illustrate that the solver’s performance both in time and memory can be greatly improved after systematic tuning of the parameters, depending on the input sparse matrix and underlying hardware.

97 MATHEMATICS AND COMPUTING↗

Direct numerical simulations for hybrid rocket boundary layers: Performance modeling and scaling

This paper presents a comprehensive performance and scaling analysis of direct numerical simulations for reacting boundary layers, focusing on slab burner configurations. Using a PETSc-based finite volume CFD framework, the study evaluates the scalability and computational cost of flow, chemistry, and radiation evaluations across 2D and 3D simulations. Polymethyl methacrylate (PMMA) is the fuel with pure O 2 as the oxidizer, modeled using a detailed chemical kinetics mechanism with 113 species and 660 reactions. A ray-tracing-based radiation solver, designed for distributed memory applications, is implemented to model radiation heat transfer. Parallel scalability is analyzed for the coupled flow, chemistry, and radiation heat transfer processes. Weak and strong scaling studies are conducted on up to 15,000 computational ranks, revealing robust performance when flow cells exceed 200 per rank. Chemistry evaluations dominate the computational cost in large 3D simulations, accounting for approximately 40% of the total runtime, while flow processes contribute around 35%, and radiation solver contributions remain below 10% due to reduced evaluation frequencies. GPU accelerated chemistry evaluation, implemented with Zero-RK, demonstrates significant promise, achieving up to a 4x speedup for workloads exceeding 30,000 cells per GPU. However, diminishing returns are observed for smaller workloads due to CPU-GPU communication overhead. This study identifies key challenges, including memory bottlenecks and the effects of domain partitioning on flow scalability, while highlighting the potential of GPU-accelerated chemistry to reduce computational costs. In conclusion, these findings provide realizable run configurations for 2D, 3D, and GPU-accelerated cases, offering insights for optimizing reactive flow solvers.

CFD Scalability↗

Microchannel-based Membrane-less Extraction of Li from Unconventional Lithium Sources & the Separation of REE

This final report provides an overview of the Project's entire duration, covering July 1, 2021 to December 31, 2023. It primarily focuses on the achievements, technological developments, and unique challenges the team faced while working on separating and extracting Lithium from produced waters. The project's primary aim was to create an integrated, high-throughput, membrane-less, and modular microfluidic platform that could extract Lithium from unconventional sources. We have successfully met all goals and milestones envisioned in the SOPO document. The most critical primary milestones, including the Go-No-Go milestone (refer to the Gantt chart in the Appendices), were successfully accomplished. We demonstrated phase separation (>90%) and extraction (>85%) performance in the MPSE using synthetic, and representative produced water composition feed at 50 ml/min total flow through MPSE 36. We have also performed a parametric study of the MPSE operations, beyond the scope of SOPO, exploring operating conditions of current and broader interest. The extended investigation of operational parameters is concurrent with our efforts to seek further development of the MPSE technology beyond the scope of the Project. Along these lines of development, we have made efforts to be responsive to DOE calls for technological developments of other types of resources (beyond PW) for the recovery of Critical Materials and higher TRL development (beyond TRL 4). During the work on this Project, we developed and implemented three innovative technical approaches that emerged from our efforts to successfully meet the Project milestones. The innovative & original technical approaches developed and implemented in this Project are now the contributions to process engineering that could be clearly credited to the Project. First, Convergent Design Approach is a comprehensive feedforward & feedback loop of four design phases: i) design for functionality, ii) design for manufacturing, iii) design for sustainability, and iv) design for market. Next was Process Intensification. A major aim of this Project was to create an innovative phase separation & extraction microscale-based technology for Li separation – thus the words microchannel-based in the Project title. A microscale-based technology is intrinsically in the center of the Process Intensification domain as defined by its unique principles. Therefore, Process Intensification was implicitly envisioned in the Project’s SOPO. Lastly, Time Scale Analysis is a novel tool for discovering the needs and directions of Process Intensification implementations in any process technology. This Project is fully credited for developing and implementing the three novel technical approaches mentioned above. These are general contributions to process engineering that emerged from this Project. Beyond the original SOPO scope, the OSU-U.Pitt research group utilized a Convergent Design methodology, integrating first-principles mathematical modeling with experimental validation on the Minimum Development Vehicle. By creating these Digital Twins, the team rapidly assessed manufacturing iterations to support TEA analysis. This framework further enabled the development of advanced Surface Modification Techniques, where hydrophobic and oleophobic coating strategies were optimized via Digital Twin tools and validated through rigorous 100-hour longevity testing. TEA Analysis: The closing efforts of this Project were focused on the TEA analysis. TEA analysis had two primary functions: i) enabling critical assessments of design variations withing 10 the Concurrent Design Approach, thus enabling evolution of the MPSE design to reach faster- better-cheaper alternatives; and ii) to create a bridge between the accomplishments of this Project and future projects of higher TRL, beyond TRL 6 level. It is important to note that the TEA model created in the Project stirred the technological solutions for the recovery of critical materials toward a vision of a very profitable modular plant that has unique zero-waste water discharge signature. More importantly, thanks to our experimental performance data and conservative assumptions, the TEA model predicts minimal technological and investment risks. Low cost of a modular unit of a nominal capacity of [1000 tons of Li 2 CO 3 /year] positions the MPSE based technology within the reach of community investors, thus offering a paradigm shift in the development of critical technologies. The project successfully navigated two primary challenges: solvent selection and manufacturing adaptation. Restricted by the SOPO to existing literature for lithium recovery, the team identified a critical need for a "material excellence program" to develop next-generation solvents, eventually concluding with a preliminary investigation into promising Ionic Liquids (ILs). Simultaneously, COVID-19 supply chain disruptions forced a pivot from traditional manufacturing to advanced additive methods at ATAMI-OSU. By transitioning from stainless steel to 3D-printed polymer substrates, the team achieved a transformative three-order-of- magnitude reduction in manufacturing costs and compressed prototyping timelines from several months to just two days. The MPSE technology offers significant energy, environmental, and economic advantages by overcoming the traditional bottlenecks of phase-separation hardware and contactor size. Unlike conventional mixer-settlers or membrane-based systems, MPSE operates without moving parts or fouling-prone membranes, achieving robust performance even with challenging, viscous, or particulate-heavy feeds. Key performance metrics include an energy intensity reduction of 5–50x (3–40 kJ/m 3 ) compared to incumbent technologies and a dramatic reduction of processing time to under 60 seconds, which drastically reduces the physical plant footprint. These technical efficiencies translate into superior economic outcomes; for a 100 t/year Li 2 CO 3 facility, implementing MPSE is projected to nearly halve contactor CAPEX (from $\$$6.08M to $\$$3.01M) and significantly increase the project's Net Present Value (NPV), derisking new investment and enabling distributed critical-mineral processing configurations. The commercialization of MPSE technology is being spearheaded by Vigsur Dynamics Inc., which has adopted a structured, parallel approach to technical and business development since its formation in January 2026. Following extensive customer discovery and engagement with the Oregon State University accelerator, Vigsur Dynamics is working to establish a business model that transitions from pilot demonstrations to modular hardware sales, ultimately aiming for a "build-own-operate" service strategy. Current technical milestones—including 100 hours of continuous operation, superior energy efficiency, and successful 6-unit modular scale-up— provide a foundation for this transition. Backed by ongoing IP licensing and a growing network of industrial and venture advisors, the company is actively de-risking the platform to replace conventional mixer-settler systems in the critical minerals market.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Automated Calibration of Parallel and Distributed Computing Simulators: A Case Study

Many parallel and distributed computing research results are obtained in simulation, using simulators that mimic real-world executions on some target system. Each such simulator is configured by picking values for parameters that define the behavior of the underlying simulation models it implements. The main concern for a simulator is accuracy: simulated behaviors should be as close as possible to those observed in the real-world target system. This requires that values for each of the simulator's parameters be carefully picked, or “calibrated,” based on ground-truth real-world executions. Examining the current state of the art shows that simulator calibration, at least in the field of parallel and distributed computing, is often undocumented (and thus perhaps often not performed) and, when documented, is described as a labor-intensive, manual process. In this work we evaluate the benefit of automating simulation calibration using simple algorithms. Specifically, we use a real-world case study from the field of High Energy Physics and compare automated calibration to calibration performed by a domain scientist. Our main finding is that automated calibration is on par with or significantly outperforms the calibration performed by the domain scientist. Furthermore, automated calibration makes it straightforward to operate desirable tradeoffs between simulation accuracy and simulation speed.

Mc donald, Jesse↗

Generating Massive Scale-free Networks: Novel Parallel Algorithms using the Preferential Attachment Model

Recently, there has been substantial interest in the study of various random networks as mathematical models of complex systems. As real-life complex systems grow larger, the ability to generate progressively large random networks becomes all the more important. This motivates the need for efficient parallel algorithms for generating such networks. Naïve parallelization of sequential algorithms for generating random networks is inefficient due to inherent dependencies among the edges and the possibility of creating duplicate (parallel) edges. In this article, we present message passing interface-based distributed memory parallel algorithms for generating random scale-free networks using the preferential-attachment model. Our algorithms are experimentally verified to scale very well to a large number of processing elements (PEs), providing near-linear speedups. The algorithms have been exercised with regard to scale and speed to generate scale-free networks with one trillion edges in 6 minutes using 1,000 PEs.

97 MATHEMATICS AND COMPUTING↗

Eliminating Yield Anisotropy and Enhancing Ductility in Mg Alloys by Shear Assisted Processing and Extrusion

Solid phase processing techniques such as friction stir welding, Shear assisted processing and extrusion (ShAPE)/ friction extrusion and cold spray have been successfully demonstrated as promising thermomechanical methods to produce metallic materials with enhanced performance. In this study, AZ series with and without silicon, ZK60 Mg alloys in as-received forms (as-cast or as-extruded) were processed using Shear Assisted Processing and Extrusion (ShAPE). Microstructural characterization was performed using EBSD and TEM and revealed that as compared to the feedstock materials/ billets, friction extruded Mg alloys had more uniform microstructure, equiaxed grains, finer and homogeneously distributed precipitates and chemical homogeneity. It was also observed that basal planes were not oriented parallel to extrusion axis. As a result, rod products exhibited significantly reduced (in some cases eliminated) yield asymmetry and achieved enhanced ductility, which were uncommon or difficult to attain using conventional processing techniques. In addition, modified texture likely suppressed deformation twinning under compressive deformation.

Shear Assisted Processing and Extrusion, SHAPE, ma↗

Development of Steady-State and Dynamic Mass and Energy Constrained Neural Networks for Distributed Chemical Systems Using Noisy Transient Data

The paper presents the development of algorithms for mass and energy constrained neural network models that can exactly conserve the overall mass and energy of distributed chemical process systems, even though the noisy transient data used for optimal model training violate the same. In contrast to approximately satisfying mass and energy balance constraints of a system by soft penalization of objective function, algorithms have been developed for solving equality-constrained nonlinear optimization problems, thus providing the guarantee of exactly satisfying the system mass and energy conservation laws. For developing dynamic mass-energy constrained network models for distributed systems, hybrid series and parallel dynamic-static neural networks have been leveraged. The developed algorithms for solving both the training and forward problems are validated using both steady-state and dynamic data in the presence of various noise characteristics. The developed data-driven algorithms are flexible to exactly satisfy mass and energy balance constraints for dynamic chemical processes if the system holdup information is available. The proposed network structures and algorithms are applied to the development of data-driven lumped and distributed models of an adiabatic superheater/reheater system, a nonisothermal continuous stirred tank reactor, as well as an electrically heated plug-flow reactor system where one form of energy gets transformed to another. It has been observed that the mass-energy constrained neural networks yield a root mean squared error of <1% with respect to the system truth for the case studies evaluated in this work.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Programming approaches for scalability, performance, and portability of combustion physics codes

Here, this paper presents the process, strategy, and results associated with porting a typical combustion physics flow solver to current state-of-the-art and future massively-parallel computer architectures. Major focus is placed on the distinct algorithmic structure of these types of codes and how it can be integrated with modern programming paradigms for heterogeneous platforms (i.e., distributed many-core systems with accelerators). An end-to-end case study is presented that exemplifies the process in a generic manner, which then serves as a clear guide with respect to the strategy and best practices leading to a robust and adaptable framework that performs well, is durable over time, is portable, and requires minimal human-effort. This end is accomplished beginning with the use of a mature, validated, structured, multiblock code framework optimized for application of both Large Eddy Simulation (LES) and Direct Numerical Simulation (DNS). This code has been ported to a variety of platforms over the past decade, including most recently the Oak Ridge Leadership Computing Facility’s “Summit” Platform. The experience gained on these multiple platforms provides general insights and thus the results presented are not specific to any one code or platform other than the overarching trend toward distributed many-core systems with accelerators in order to move toward exascale performance. The resultant performance and scalability of the ported code is demonstrated on a real-world application; a state-of-the-art rotating detonation rocket engine simulation that matches the complex geometry and boundary conditions imposed as part of a companion experimental campaign.

97 MATHEMATICS AND COMPUTING↗

Scalable training of trustworthy and energy-efficient predictive graph foundation models for atomistic materials modeling: a case study with HydraGNN

We present our work on developing and training scalable, trustworthy, and energy-efficient predictive graph foundation models (GFMs) using HydraGNN, a multi-headed graph convolutional neural network architecture. HydraGNN expands the boundaries of graph neural network (GNN) computations in both training scale and data diversity. It abstracts over message passing algorithms, allowing both reproduction of and comparison across algorithmic innovations that define nearest-neighbor convolution in GNNs. This work discusses a series of optimizations that have allowed scaling up the GFMs training to tens of thousands of GPUs on datasets consisting of hundreds of millions of graphs. Our GFMs use multitask learning (MTL) to simultaneously learn graph-level and node-level properties of atomistic structures, such as energy and atomic forces. Using over 154 million atomistic structures for training, we illustrate the performance of our approach along with the lessons learned on two state-of-the-art US Department of Energy (US-DOE) supercomputers, namely the Perlmutter petascale system at the National Energy Research Scientific Computing Center and the Frontier exascale system at Oak Ridge Leadership Computing Facility. The HydraGNN architecture enables the GFM to achieve near-linear strong scaling performance using more than 2000 GPUs on Perlmutter and 16,000 GPUs on Frontier.

97 MATHEMATICS AND COMPUTING↗

On-chip parallel processing of quantum frequency comb

Abstract The frequency degree of freedom of optical photons has been recently explored for efficient quantum information processing. Significant reduction in hardware resources and enhancement of quantum functions can be expected by leveraging the large number of frequency modes. Here, we develope an integrated photonic platform for the generation and parallel processing of quantum frequency combs (QFCs). Cavity-enhanced parametric down-conversion with Sagnac configuration is implemented to generate QFCs with identical spectral distributions. On-chip quantum interference of different frequency modes is simultaneously realized with the same photonic circuit. High interference visibility is maintained across all frequency modes with the identical circuit setting. This enables the on-chip reconfiguration of QFCs. By deterministically separating QFCs without spectral filtering, we further demonstrate high-dimensional Hong-Ou-Mandel effect. Our work provides the critical step for the efficient implementation of quantum information processing with integrated photonics using the frequency degree of freedom.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

UPC++ v1.0 Programmer’s Guide (Rev. 2023.9.0)

UPC++ is a C++ library that supports Partitioned Global Address Space (PGAS) programming. It is designed for writing efficient, scalable parallel programs on distributed-memory parallel computers. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). The UPC++ control model is single program, multiple-data (SPMD), with each separate constituent process having access to local memory as it would in C++. The PGAS memory model additionally provides one-sided RMA communication to a global address space, which is allocated in shared segments that are distributed over the processes. UPC++ also features Remote Procedure Call (RPC) communication, making it easy to move computation to operate on data that resides on remote processes. UPC++ was designed to support exascale high-performance computing, and the library interfaces and implementation are focused on maximizing scalability. In UPC++, all communication operations are syntactically explicit, which encourages programmers to consider the costs associated with communication and data movement. Moreover, all communication operations are asynchronous by default, encouraging programmers to seek opportunities for overlapping communication latencies with other useful work. UPC++ provides expressive and composable abstractions designed for efficiently managing aggressive use of asynchrony in programs. Together, these design principles are intended to enable programmers to write applications using UPC++ that perform well even on hundreds of thousands of cores.

97 MATHEMATICS AND COMPUTING↗

UPC++ v1.0 Programmer’s Guide (Revision 2022.3.0)

UPC++ is a C++ library that supports Partitioned Global Address Space (PGAS) programming. It is designed for writing efficient, scalable parallel programs on distributed-memory parallel computers. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). The UPC++ control model is single program, multiple-data (SPMD), with each separate constituent process having access to local memory as it would in C++. The PGAS memory model additionally provides one-sided RMA communication to a global address space, which is allocated in shared segments that are distributed over the processes. UPC++ also features Remote Procedure Call (RPC) communication, making it easy to move computation to operate on data that resides on remote processes. UPC++ was designed to support exascale high-performance computing, and the library interfaces and implementation are focused on maximizing scalability. In UPC++, all communication operations are syntactically explicit, which encourages programmers to consider the costs associated with communication and data movement. Moreover, all communication operations are asynchronous by default, encouraging programmers to seek opportunities for overlapping communication latencies with other useful work. UPC++ provides expressive and composable abstractions designed for efficiently managing aggressive use of asynchrony in programs. Together, these design principles are intended to enable programmers to write applications using UPC++ that perform well even on hundreds of thousands of cores.

97 MATHEMATICS AND COMPUTING↗

UPC++ v1.0 Programmer’s Guide, Revision 2023.3.0

UPC++ is a C++ library that supports Partitioned Global Address Space (PGAS) programming. It is designed for writing efficient, scalable parallel programs on distributed-memory parallel computers. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). The UPC++ control model is single program, multiple-data (SPMD), with each separate constituent process having access to local memory as it would in C++. The PGAS memory model additionally provides one-sided RMA communication to a global address space, which is allocated in shared segments that are distributed over the processes. UPC++ also features Remote Procedure Call (RPC) communication, making it easy to move computation to operate on data that resides on remote processes. UPC++ was designed to support exascale high-performance computing, and the library interfaces and implementation are focused on maximizing scalability. In UPC++, all communication operations are syntactically explicit, which encourages programmers to consider the costs associated with communication and data movement. Moreover, all communication operations are asynchronous by default, encouraging programmers to seek opportunities for overlapping communication latencies with other useful work. UPC++ provides expressive and composable abstractions designed for efficiently managing aggressive use of asynchrony in programs. Together, these design principles are intended to enable programmers to write applications using UPC++ that perform well even on hundreds of thousands of cores.

97 MATHEMATICS AND COMPUTING↗

UPC++ v1.0 Programmer’s Guide, Revision 2022.9.0

UPC++ is a C++ library that supports Partitioned Global Address Space (PGAS) programming. It is designed for writing efficient, scalable parallel programs on distributed-memory parallel computers. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). The UPC++ control model is single program, multiple-data (SPMD), with each separate constituent process having access to local memory as it would in C++. The PGAS memory model additionally provides one-sided RMA communication to a global address space, which is allocated in shared segments that are distributed over the processes. UPC++ also features Remote Procedure Call (RPC) communication, making it easy to move computation to operate on data that resides on remote processes. UPC++ was designed to support exascale high-performance computing, and the library interfaces and implementation are focused on maximizing scalability. In UPC++, all communication operations are syntactically explicit, which encourages programmers to consider the costs associated with communication and data movement. Moreover, all communication operations are asynchronous by default, encouraging programmers to seek opportunities for overlapping communication latencies with other useful work. UPC++ provides expressive and composable abstractions designed for efficiently managing aggressive use of asynchrony in programs. Together, these design principles are intended to enable programmers to write applications using UPC++ that perform well even on hundreds of thousands of cores.

97 MATHEMATICS AND COMPUTING↗

Leveraging Pre-Built Catalogs and Object-Level Scheduling to Eliminate I/O Bottlenecks in HPC Environments

Modern High-Performance Computing (HPC) environments face mounting challenges due to the shift from large to small file datasets, along with an increasing number of users and parallelized applications. As HPC systems rely on Parallel File Systems (PFS), such as Lustre for data processing, performance bottlenecks stemming from Object Storage Target (OST) contention have become a significant concern. Existing solutions, such as LADS with its object-level scheduling approach, fall short in large-scale HPC environments due to their inability to effectively address metadata I/O bottlenecks and the growing number of I/O processes. This study highlights the pressing need for a comprehensive solution that tackles both OST contention and metadata I/O challenges in diverse HPC workloads. To address these challenges, we propose SwiftLoad, an object-level I/O scheduling framework that leverages a metadata catalog to enhance the performance and efficiency of parallel HPC utilities. The adoption of the metadata catalog mitigates the metadata I/O bottlenecks that commonly occur in HPC utilities, a challenge that is particularly pronounced in object-level I/O scheduling. SwiftLoad addresses OST contention and the uneven distribution of I/O processes across different OSTs through mathematical modeling and incorporates a Loader Configuration Module to regulate the number of I/O processes. Evaluated with two representative utilities—data deduplication profiling and data augmentation—SwiftLoad achieved performance improvements of up to 5.63x and 11.0x, respectively, on a production supercomputer.

HPC↗

Proximity Portability and in Transit , M-to-N Data Partitioning and Movement in SENSEI [Book Chapter]

In high-performance parallel in situ processing, the term in transit processing refers to those configurations where data must move from a producer to a consumer that runs on separate resources. In the context of parallel and distributed computing on an HPC platform one of the central challenges is to determine a mapping of data from producer ranks to consumer ranks. This problem is complicated by the heterogeneity that arises in producer-consumer pairs, such as when producer and consumer codes have different levels of concurrency, different scaling characteristics, or different data models. The resulting mapping and movement of data from M producer to N consumer ranks can have a significant impact on aggregate application performance, particularly when the data consumer requires only a subset of the overall data for its task. This chapter focuses on the design considerations that underlie SENSEI’s implementation to this challenging problem. These design considerations extend the core SENSEI architecture and include ideas like the need to accommodate flexibility in the choice of different partitioning methods, the ability for a data consumer to request and receive only the subset of data needed for its particular operation, and the ability to leverage any of several different data transport tools. The idea of proximity portability, being able to use different data transport methods as part of an in transit workflow, is illustrated through the use of three different transport layers where switching from one transport tool to another is accomplished with only a configuration file change. Here, the chapter also includes a performance analysis summary showing the performance gains that are possible in terms of multiple metrics, such as memory footprint, time to solution, and amount of data moved, when using optimized partitioners in an in transit setting, gains that are made possible by the implementation shaped by specific design considerations.

Bethel, E. Wes↗

Sequence length scaling in vision transformers for scientific images on frontier

Vision Transformers (ViTs) are pivotal for foundational models in scientific imagery, including Earth science applications, due to their capability to process large sequence lengths. While transformers for text have inspired scaling sequence lengths in ViTs, adapting these for ViTs introduces unique challenges. We develop distributed sequence parallelism for ViTs, enabling them to handle up to 1M tokens. Our approach, leveraging DeepSpeed-Ulysses and Long-Sequence-Segmentation with model sharding, is the first to apply sequence parallelism in ViT training, achieving a 94% batch scaling efficiency on 2,048 AMD-MI250X GPUs. Evaluating sequence parallelism in ViTs, particularly in models up to 10B parameters, highlighted substantial bottlenecks. We countered these with hybrid sequence, pipeline, and flash attention strategies, to scale beyond single GPU memory limits. Our method significantly enhances climate modeling accuracy by 20% in temperature predictions, marking the first training of a vision transformer model to convergence with a sequence length of 188K tokens, using full self-attention.

Tsaris, Aristeidis (aris) [ORNL] (ORCID:0000000277↗

Current Practices in Distribution Utility Resilience Planning for Winter Storms

This report is part of a series of hazard-focused case studies examining common practices in electric utility resilience planning. We use standard terminology defining resilience as the ability to anticipate, withstand, absorb, and recover from hazards that cause long duration outages. We distinguish between reliability and resilience using Institute of Electrical and Electronics Engineers (IEEE) 1366-2022, which defines major events as an event that exceeds reasonable design and/or operational limits of the electric power system. Resilience planning is focused on major event days and reliability planning is focused on nonmajor event days. Utility resilience plans are assessed according to common resilience components identified in existing resilience frameworks. The focus of this report is on winter storms in which the primary hazards are heavy snowfall, freezing rain, ice, extreme cold, severe wind, and flooding. These hazards can also contribute to generation shortages, resulting in bulk power system impacts that have consequences for the distribution system, such as load shedding. Stand-alone reports focusing on wildfires and nonwinter storms have been published in parallel with this report. This report can be used as a starting point for understanding potential investment prioritization processes and investment options. This report is intended to improve utility resilience planning by supporting constructive dialogue among utilities, regulators, and other stakeholders.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗