Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “load balancing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Photovoltaic Analysis and Response Support (PARS) Platform for Solar Situational Awareness and Resiliency Services

The project's primary objective is to develop a digital-twin based Photovoltaic (PV) Analysis and Response Support (PARS) platform, which aims to provide real-time situational awareness and optimal response plans. This platform is designed to enhance the performance of hybrid PV systems, making them competitive with or even superior to conventional generation resources. The PARS platform enabled the project team to develop and evaluate an extensive suite of grid support functionalities for the hybrid PV systems to enhance grid performance, across key areas including visibility, dispatchability, security, resilience, and reliability. Given the global push toward achieving 100% clean energy by 2035, there is a significant increase in the integration of inverter-based resources (IBRs) throughout the energy grid. Effectively managing the inherent variability and uncertainty associated with IBRs is crucial for ensuring cost-effectiveness, reliability, and security in both the main grid and islanded microgrids. Constrained to a limited array of IEEE test systems or standard feeder models, traditional IBR modeling struggles to assimilate new field data, accurately reflect system dynamics, and adapt to the evolving energy landscape. In our project, we embraced a Digital Twin (DT) strategy for crafting the PARS platform. A digital twin acts as a precise virtual counterpart of a physical system, built on historical data and continuously honed with real-time insights. This enables the high-fidelity DT to accurately mirror current system operations and forecast future scenarios. Consequently, the PARS platform becomes an ideal environment for testing and refining monitoring, control, power, and energy management algorithms designed to boost hybrid PV system performance. The defining feature of the PARS platform, distinguishing it from other advanced simulation tools, is its exceptional adaptability. This is achieved by employing actual network topologies and utilizing real-time field data for fine-tuning and calibration, ensuring a close emulation of real-world conditions. The project deliverables include: 1) High-fidelity IBR models and tools for real-time parameterization, utilizing real-time field measurements to refine IBR models for enhanced accuracy and performance; 2) Grid-forming and Grid-following capabilities to deliver resilience services, including blackstart, voltage and frequency support, cold-load pick-up, power reserves, and three-phase load balancing across grid-connected and microgrid settings; 3) Machine learning-based forecasting tools and methods for generating synthetic data and topologies, creating diverse and realistic simulation environments for evaluating varied operational scenarios; 4) Advanced microgrid power and energy management algorithms for optimizing the integration and operation of PV, storage, and demand response resources within both feeder and community scales. The power grid data sets are provided by four utility companies in North Carolina and the New York Power Administration. Acting as industry advisors, our industry partners communicated stakeholder needs and regulatory standards to the research teams, aiding technology transfer by incorporating the developed methodologies into their daily operations. This collaboration ensures that the PARS platform, functioning as a power system digital twin, enhances our understanding of IBR dynamic behaviors and enables the development and evaluation of IBR control functions that match or exceed the capabilities of conventional synchronous generators.

14 SOLAR ENERGY↗

Runtime performance of a GAMESS quantum chemistry application offloaded to GPUs

Summary Computational chemistry is at the forefront of solving urgent societal problems, such as polymer upcycling and carbon capture. The complexity of modeling these processes at appropriate length and time scales is mainly manifested in the number and types of chemical species involved in the reactions and may require models of several thousand atoms and large basis sets to accurately capture the chemical complexity and heterogeneity in the physical and chemical processes. The quantum chemistry package General Atomic and Molecular Electronic Structure System (GAMESS) has a wide array of methods that can efficiently and accurately treat complex chemical systems. In this work, we have used the GAMESS Effective Fragment Molecule Orbital (EFMO) method for electronic structure calculation of a challenging mesoporous silica nanoparticle (MSN) model surrounded by about 4700 water molecules to investigate the strong scaling and GPU offloading on hybrid CPU‐GPU nodes. Experiments were performed on the Perlmutter platform at the National Energy Research Scientific Computing Center. Good strong scaling and load balancing have been observed on up to 88 hybrid nodes for different settings of the execution parameters for the calculation considered here. When GPUs are oversubscribed by offloading work from multiple CPU processes, using the NVIDIA multi‐process service (MPS) has consistently reduced time to solution and energy consumed. Additionally, for some configuration parameter settings, oversubscription with MPS improved performance by up to 5.8% over the case without oversubscription.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Energy-efficient cooperative resource allocation and task scheduling for Internet of Things environments

Offloading Internet of Things (IoT) tasks to the cloud for further processing might not always lead to an optimal execution time, particularly in situations such as resource contention, under-provisioning, over-provisioning, and fragmentation. In addition, dynamically optimizing the number of Virtual Machines (VMs) for resource scheduling in order to meet application requirements remains a major research challenge. Further, existing resource scheduling algorithms focus primarily on minimizing operational costs while maximizing resource sharing and utilization. Considering energy utilization as part of the resource allocation and scheduling process as an optimization objective for maintaining load balancing has often been neglected. To address these challenges and more, we propose a cooperative energy-aware resource allocation and scheduling strategy based on a Technique for Order of Preference by Similarity to Ideal Solution (TOPSIS) multi-criteria decision-making method. Here we used the Grid Workloads Archive dataset to evaluate our proposed approach named TOPREAL. Experimental results with respect to the allocation of VM resources when considering processing a large segment of tasks indicate that TOPREAL outperforms existing algorithms in terms of energy savings, with an average improvement of 40.25%, while maintaining an average improvement of 16.21% when it comes to execution time. Results also demonstrate that our method can save an average of 78.06 processing hours and 63,215kJ of energy when compared to existing scheduling algorithms. These results demonstrate the effectiveness of our proposed model and the viability of using multi-criteria decision-making techniques such as TOPSIS to solve the resource allocation and scheduling problem in edge environments.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

An adaptive scalable fully implicit algorithm based on stabilized finite element for reduced visco-resistive MHD

The magnetohydrodynamics (MHD) equations are continuum models used in the study of a wide range of plasma physics systems, including the evolution of complex plasma dynamics in tokamak disruptions. However, efficient numerical solution methods for MHD are extremely challenging due to disparate time and length scales, strong hyperbolic phenomena, and nonlinearity. Additionally, therefore the development of scalable, implicit MHD algorithms and high-resolution adaptive mesh refinement strategies is of considerable importance. In this work, we develop a high-order stabilized finite-element algorithm for the reduced visco-resistive MHD equations based on the MFEM finite element library (mfem.org). The scheme is fully implicit, solved with the Jacobian-free Newton-Krylov (JFNK) method with a physics-based preconditioning strategy. Our preconditioning strategy is a generalization of the physics-based preconditioning methods in Chacón et al. (2002) to adaptive, stabilized finite elements. Algebraic multigrid methods are used to invert sub-block operators to achieve scalability. A parallel adaptive mesh refinement scheme with dynamic load-balancing is implemented to efficiently resolve the multi-scale spatial features of the system. Our implementation uses the MFEM framework, which provides arbitrary-order polynomials and flexible adaptive conforming and non-conforming meshes capabilities. Results demonstrate the accuracy, efficiency, and scalability of the implicit scheme in the presence of large scale disparity. The potential of the AMR approach is demonstrated on an island coalescence problem in the high Lundquist-number regime (≥ 10 7 ) with the successful resolution of plasmoid instabilities and thin current sheets.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Scalable Implicit Solvers with Dynamic Mesh Adaptation for a Relativistic Drift-Kinetic Fokker–Planck–Boltzmann Model

In this work we consider a relativistic drift-kinetic model for runaway electrons along with a Fokker–Planck operator for small-angle Coulomb collisions, a radiation damping operator, and a secondary knock-on (Boltzmann) collision source. Here, we develop a new scalable fully implicit solver utilizing finite volume and conservative finite difference schemes and dynamic mesh adaptivity. A new data management framework in the PETSc library based on the p4est library is developed to enable simulations with dynamic adaptive mesh refinement (AMR), distributed memory parallelization, and dynamic load balancing of computational work. This framework and the runaway electron solver building on the framework are able to dynamically capture both bulk Maxwellian at the low-energy region and a runaway tail at the high-energy region. To effectively capture features via the AMR algorithm, a new AMR indicator prediction strategy is proposed that is performed alongside the implicit time evolution of the solution. This strategy is complemented by the introduction of computationally cheap feature-based AMR indicators that are analyzed theoretically. Numerical results quantify the advantages of the prediction strategy in better capturing features compared with nonpredictive strategies; and we demonstrate trade-offs regarding computational costs. The robustness with respect to model parameters, algorithmic scalability, and parallel scalability are demonstrated through several benchmark problems including manufactured solutions and solutions of different physics models. We focus on demonstrating the advantages of using implicit time stepping and AMR for runaway electron simulations.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A cell-centered AMR-ALE framework for 3D multi-material hydrodynamics. Part I: Lagrangian and indirect Euler AMR algorithms

Many applications of physics and engineering involve wide ranges of time and spatial scales. The numerical simulation of localized small scales such as shock waves and material interfaces requires a large number of computational cells in these regions. For these applications, Lagrangian and Arbitrary-Lagrangian-Eulerian (ALE) related methods are engaging since the moving mesh feature naturally brings mesh cells on shock discontinuities and material interfaces are carefully captured. In addition, Adaptive-Mesh-Refinement (AMR) strategies aim to optimize computational resources by concentrating finer mesh cells only in areas of interest while using coarser cells elsewhere. A key but challenging AMR requirement consists in efficiently distributing the computational effort to achieve high accuracy without the prohibitive computational costs associated with uniformly fine grids. Here, in this document, the coupling of the p4est AMR library with a cell-centered Lagrangian scheme is presented with the goal to perform reliable 3D Lagrangian-AMR and indirect Euler-AMR multi-material simulations. In particular, it is shown that starting from a 3D indirect ALE code, the memory management and load balancing requirements can be delegated to an external library (here the p4est library) to unlock ALE-AMR capabilities. First, we present a strategy to transcribe the octant-based connectivity of the 3D AMR framework with that of an unstructured mesh of polygonal cells used in Lagrangian hydrodynamics. Then, we show how refinement and coarsening operations must be adapted to the particular Lagrangian framework to ensure the conservation of volume during those steps. Finally, several numerical test cases are presented that demonstrate the capabilities of the Lagrangian-AMR and indirect Euler-AMR algorithms.

3D cell-centered Lagrangian numerical scheme↗

Towards scaling community detection on distributed-memory heterogeneous systems

Distributed multi-GPU systems pose significant challenges and opportunities for efficient execution of parallel applications. Graph algorithms are generally characterized by irregular memory accesses, low computation to communication ratios, and load balancing problems that are especially hard to address on multi-GPU systems. Graph community detection is an important problem in the emerging domain of graph analytics with numerous applications. In this paper, we present our ongoing work on distributed-memory multi-GPU implementation for graph community detection. Our work parallelizes the widely used (albeit serial) Louvain method on distributed multi-GPU platforms. Supported by an extensive set of experiments on a multi-GPU enabled supercomputer (OLCF Summit) and a single compute node (Nvidia DGX-2®), we demonstrate competitive performance to existing distributed-memory CPU-based implementation, and up to 6.5 better results than Nvidia RAPIDS® CUGRAPH. To the best of our knowledge, this work represents the first effort for community detection on distributed multi-GPU systems. Our approach and related findings can be extended to numerous other iterative graph algorithms on multi-GPU systems.

97 MATHEMATICS AND COMPUTING↗

Acceleration of the particle-in-cell code Osiris with graphics processing units

Fully relativistic particle-in-cell (PIC) simulations are crucial for advancing our knowledge of plasma physics. Modern supercomputers based on graphics processing units (GPUs) offer the potential to perform PIC simulations of unprecedented scale, but require robust and feature-rich codes that can fully leverage their computational resources. In this work, this demand is addressed by adding GPU acceleration to the PIC code Osiris. An overview of the algorithm, which features a CUDA extension to the underlying Fortran architecture, is given. Detailed performance benchmarks for thermal plasmas are presented, which demonstrate excellent weak scaling on NERSC's Perlmutter supercomputer and high levels of absolute performance. The robustness of the code to model a variety of physical systems is demonstrated via simulations of Weibel filamentation and laser-wakefield acceleration run with dynamic load balancing. Finally, measurements and analysis of energy consumption are provided that indicate that the GPU algorithm is up to ~14 times faster and ~7 times more energy efficient than the optimized CPU algorithm on a node-to-node basis. The described development addresses the PIC simulation community's computational demands both by contributing a robust and performant GPU-accelerated PIC code and by providing insight into efficient use of GPU hardware.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

BEYONDPLANCK III. Commander3

We describe the computational infrastructure for end-to-end Bayesian cosmic microwave background (CMB) analysis implemented by the BeyondPlanck Collaboration. The code is called Commander3. It provides a statistically consistent framework for global analysis of CMB and microwave observations and may be useful for a wide range of legacy, current, and future experiments. The paper has three main goals. Firstly, we provide a high-level overview of the existing code base, aiming to guide readers who wish to extend and adapt the code according to their own needs or re-implement it from scratch in a different programming language. Secondly, we discuss some critical computational challenges that arise within any global CMB analysis framework, for instance in-memory compression of time-ordered data, fast Fourier transform optimization, and parallelization and load-balancing. Thirdly, we quantify the CPU and RAM requirements for the current BEYONDPLANCK analysis, finding that a total of 1.5 TB of RAM is required for efficient analysis and that the total cost of a full Gibbs sample for LFI is 170 CPU-hrs, including both low-level processing and high-level component separation, which is well within the capabilities of current low-cost computing facilities. The existing code base is made publicly available under a GNU General Public Library (GPL) license.

79 ASTRONOMY AND ASTROPHYSICS↗

Federated Access from DOE Labs to Distributed Storage in the EIC Era of Computing

The Electron Ion Collider (EIC) collaboration and future experiment is a unique scientific ecosystem within Nuclear Physics as the experiment starts right off as a crosscollaboration between Brookhaven National Lab (BNL) & Jefferson Lab (JLab). As a result, this muti-lab computing model tries at best to provide services accessible from anywhere by anyone who is part of the collaboration. While the computing model for the EIC is not finalized, it is anticipated that the computational and storage resources will be made accessible to a wide range of collaborators across the world. The use of federated ID seems to be a critical element to the strategy of providing such services, allowing seamless access to each lab site computing resources. However, providing Federated access to a Federated storage is not a trivial matter and has its share of technical challenges. In this contribution, we focus on the steps we took towards the deployment of a distributed object storage system that integrates with Amazon S3 and Federated ID. We will first cover for and explain the first stage storage solutions provided to the EIC during the detector design phase. Our initial test deployment consisted of Lustre storage using MinIO, hence providing an S3 interface. High Availability load balancers were added later to provide the initial scalability it lacked. Performance of that system will be shown. While this embryonic solution worked well, it had many limitations. Looking ahead, the Ceph object storage is considered a top-of-the-line solution in the storage community - since the Ceph Object Gateway is compatible with the Amazon S3 API out of the box, our next phase will use a native S3 storage. Our Ceph deployment will consist of erasure coded storage nodes to maximize storage potential along with multiple Ceph Object Gateways for redundant access. We will compare performance of our next stage implementations. Finally, we will present how to leverage OpenID Connect with the Ceph Object Gateway’s to enable Federated ID access. We hope this contribution will serve the community needs as we move forward with cross-lab collaborations and the need for Federated ID access to distributed compute facilities.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Multi-level parallelization of quantum-chemical calculations

Here, strategies for multiple-level parallelizations of quantum-mechanical calculations are discussed, with an emphasis on using groups of workers for performing parallel tasks. These parallel programming models can be used for a variety ab initio quantum chemistry approaches, including the fragment molecular orbital method and replica-exchange molecular dynamics. Strategies for efficient load balancing on problems of increasing granularity are introduced and discussed. A four-level parallelization is developed based on a multi-level hierarchical grouping, and a high parallel efficiency is achieved on the Theta supercomputer using 131 072 OpenMP threads.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A parallel, distributed memory implementation of the adaptive sampling configuration interaction method

The many-body simulation of quantum systems is an active field of research that involves several different methods targeting various computing platforms. Many methods commonly employed, particularly coupled cluster methods, have been adapted to leverage the latest advances in modern high-performance computing. Selected configuration interaction (sCI) methods have seen extensive usage and development in recent years. However, the development of sCI methods targeting massively parallel resources has been explored only in a few research works. Here, we present a parallel, distributed memory implementation of the adaptive sampling configuration interaction approach (ASCI) for sCI. In particular, we will address the key concerns pertaining to the parallelization of the determinant search and selection, Hamiltonian formation, and the variational eigenvalue calculation for the ASCI method. Load balancing in the search step is achieved through the application of memory-efficient determinant constraints originally developed for the ASCI-PT2 method. The presented benchmarks demonstrate near optimal speedup for ASCI calculations of Cr 2 (24e, 30o) with 10 6 , 10 7 , and 3 × 10 8 variational determinants on up to 16 384 CPUs. Importantly, to the best of the authors’ knowledge, this is the largest variational ASCI calculation to date.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Cryogenic aspects of a 20 MW class low-temperature superconducting generator for the renewables industry

In this paper we give a progress update on the cryogenic design for cooling a 20 MW class, partially superconducting generator with stationary field coils for offshore wind renewables industry [1-2]. This is a continuation of an earlier program on a 10 MW system dating back ten years. Whereas this new power rating increase leads to a radial diameter expansion of the superconducting field coils from 4 to > 9.5 m, it maintains the axial length. We show the design based on this scaled up with enlarged diameter. The field coil size increase asks for higher cooling power that leads to a bigger cold box size to accommodate the cryocoolers. In the design, as many as 8 cryocoolers can be accessed and serviced from the nacelle that houses the main cryogenic components. A typical thermal load balance sheet with all components is given and compared against the available cryocooler cooling power at different operating conditions. Due to the cryocoolers’ local point of contact cooling within the nacelle, the temperature gradient of the thermal shield that fully encloses the cold mass with its embedded field coils needs to be balanced out to minimize the heat burden on the field coils. This requires additional analytical efforts and implementation of further design features. The extended nacelle houses the cold box with its cryogenic infrastructure. The interface for the envisaged cryogenic pushbutton closed-loop circulating system remains invisible, requires no handling of cryogenic liquids and is hermetically closed. The field coil diameter increase also leads to a greater initial helium gas storage volume. In addition, it requires higher initial room temperature fill pressure within the toroidal helium vapor storage tanks and in cooling tubes connecting to the cryocoolers. The toroidal helium gas tanks are thermally coupled with the thermal shield so that helium convection inside the storage tank can improve heat transfer and reduce the temperature gradient of the thermal shield. Besides those heat transfer challenges, additional mechanical strain within this large structure is exerted on the torque tubes during initial cooldown and when energizing field coils. Some of those design challenges are quite unexpected, leading to novel workarounds in order to maintain the chosen cooling strategy. Finally, we assess those design limitations in view of further cryogenic scalability with emphasis on manufacturability and assembly.

cryogenics↗

From Cell to System: Accelerated hpc Simulations of BESS Aging under Frequency Regulation and Arbitrage use cases

Lithium-ion battery energy storage systems (BESS) packs have emerged as a leading solution for grid-scale energy storage, enhancing resiliency and balancing load fluctuations. Yet, experimental characterization of large-format LIB packs-particularly to assess performance and degradation over hundreds of cycles - demands substantial hardware investment and multi-year testing campaigns. In this work, we couple a hierarchical, physics-based modeling framework agnostic to electrode chemistries with high-performance computing to accelerate systems level evaluation by upto two orders of magnitude. Building on the open-source liionpack platform, we implement cell, module, and pack-scale electrochemical models enriched with mechanistic aging mechanisms and deploy them on an HPC cluster to simulate 150−200kWh systems over 500 - 1,000 cycles with in days. We subject these virtual B ESS to both constant-current cycling and realistic grid service profiles spanning frequency regulation, ramp-rate support, and energy arbitrage-and quantify the resulting degradation patterns. Our results reveal that localized cell aging can induce substantial nonuniformity at module and pack levels, with service-specific cycling protocols driving distinct aging modes. This rapid, multiscale modeling approach provides a powerful design-space exploration tool for optimizing electrical architecture, control strategies, and operational schedules to prolong pack lifetime and lower total cost of ownership.

Ayalasomayajula, Surya [ORNL] (ORCID:0009000860788↗

MICCO: An Enhanced Multi-GPU Scheduling Framework for Many-Body Correlation Functions

Calculation of many-body correlation functions is one of the critical kernels utilized in many scientific computing areas, especially in Lattice Quantum Chromodynamics (Lattice QCD). It is formalized as a sum of a large number of contraction terms each of which can be represented by a graph consisting of vertices describing quarks inside a hadron node and edges designating quark propagations at specific time intervals. Due to its computation- and memory-intensive nature, real-world physics systems (e.g., multi-meson or multi-baryon systems) explored by Lattice QCD prefer to leverage multi-GPUs. Different from general graph processing, many-body correlation function calculations show two specific features: a large number of computation-/data-intensive kernels and frequently repeated appearances of original and intermediate data. The former results in expensive memory operations such as tensor movements and evictions. The latter offers data reuse opportunities to mitigate the data-intensive nature of many-body correlation function calculations. However, existing graph-based multi-GPU schedulers cannot capture these data-centric features, thus resulting in a sub-optimal performance for many-body correlation function calculations. To address this issue, this paper presents a multi-GPU scheduling framework, MICCO, to accelerate contractions for correlation functions particularly by taking the data dimension (e.g., data reuse and data eviction) into account. This work first performs a comprehensive study on the interplay of data reuse and load balance, and designs two new concepts: local reuse pattern and reuse bound to study the opportunity of achieving the optimal trade-off between them. Based on this study, MICCO proposes a heuristic scheduling algorithm and a machine-learning-based regression model to generate the optimal setting of reuse bounds. Specifically, MICCO is integrated into a real-world Lattice QCD system, Redstar, for the first time running on multiple GPUs. The evaluation demonstrates MICCO outperforms other state-of-art works, achieving up to 2.25× speedup in synthesized datasets, and 1.49× speedup in real-world correlation functions.

Wang, Qihan↗

Tiling Framework for Heterogeneous Computing of Matrix based Tiled Algorithms

Tiling matrix operations can improve the load balancing and performance of applications on heterogeneous computing resources. Writing a tile-based algorithm for each operation with a traditional, hand-tuned tiling approach that uses for loops in C/C++ is cumbersome and error prone. Moreover, it must enable and support the heterogeneous memory management of data objects and also explore architecture-supported, native, tiled-data transfer APIs instead of copying the tiled data to continuous memory before the data transfer. The tiling framework provides a tiled data structure for heterogeneous memory mapping and parameterization to a heterogeneous task specification API. We have integrated our tiled framework into MatRIS (Math kernels library using IRIS). IRIS is a heterogeneous run-time framework with a heterogeneous programming model, memory model, and task execution model. Experiments reveal that the tiled framework for BLAS operations has improved the programmability of tiled BLAS and improved performance by ~20% when compared against the traditional method that copies the data to continuous memory locations for heterogeneous computing.

Miniskar, Narasinga Rao↗

RingX: Scalable Parallel Attention for Long-Context Learning on HPC

The attention mechanism has become foundational for remarkable AI breakthroughs since the introduction of the Transformer, driving the demand for increasingly longer context to power frontier models such as large-scale reasoning language models and high-resolution image/video generators. However, its quadratic computational and memory complexities present substantial challenges. Current state-of-the-art parallel attention methods, such as ring attention, are widely adopted for long-context training but utilize a point-to-point communication strategy that fails to fully exploit the capabilities of modern HPC network architectures. In this work, we propose ringX, a scalable family of parallel attention methods optimized explicitly for HPC systems. By enhancing workload partitioning, refining communication patterns, and improving load balancing, ringX achieves up to 3.4 × speedup compared to conventional ring attention on the Frontier supercomputer. Optimized for both bi-directional and causal attention mechanisms, ringX demonstrates its effectiveness through training benchmarks of a Vision Transformer (ViT) on a climate dataset and a Generative Pre-Trained Transformer (GPT) model, Llama3 8B. Our method attains an end-to-end training speedup of approximately 1.5 × in both scenarios. To our knowledge, the achieved 38% model FLOPs utilization (MFU) for training Llama3 8B with a 1M-token sequence length on 4,096 GPUs represents one of the highest training efficiencies reported for long-context learning on HPC systems. Our code implementation is available at https://github.com/jqyin/ringX-attention.

Yin, Junqi [ORNL] (ORCID:0000000338435520)↗

Xyce(™) Parallel Electronic Simulator v.7.5

The Xyce Parallel Electronic Simulator simulates electronic circuit behavior in DC, AC, HB, MPDE and transient mode using standard analog (DAE) and/or device (PDE) device models including several age and radiation aware devices. It supports a variety of computing platforms (both serial and parallel) computers. Lastly, it uses a variety of modern solution algorithms dynamic parallel load-balancing and iterative solvers.! ! Xyce is primarily used to simulate the voltage and current behavior of a circuit network (a network of electronic devices connected via a conductive network). As a tool, it is mainly used for the design and analysis of electronic circuits.! ! Kirchoff's conservation laws are enforced over a network using modified nodal analysis. This results in a set of differential algebraic equations (DAEs). The resulting nonlinear problem is solved iteratively using a fully coupled Newton method, which in turn results in a linear system that is solved by either a standard sparse-direct solver or iteratively using Trilinos linear solver packages, also developed at Sandia National Laboratories.

Source record↗