Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel machines”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Propulsion Electrification Architecture Selection Process and Cost of Carbon Abatement Analysis for Heavy-Duty Off-Road Material Handler

The heavy-duty off-road industry continues to expand efforts to reduce fuel consumption and CO 2 e (carbon dioxide equivalent) emissions. Many manufacturers are pursuing electrification to decrease fuel consumption and emissions. Future policies will likely require electrification for CO2e savings, as seen in light-duty on-road vehicles. Electrified architectures vary widely in the heavy-duty off-road space, with parallel hybrids in some applications and series hybrids in others. The diverse applications for different types of equipment mean different electrified configurations are required. Companies must also determine the value in pursuing electrified architectures; this work analyzes a range of electrified architectures, from micro hybrids to parallel hybrids to series hybrids to a BEV, looking at the total cost, total CO 2 e, and cost per CO 2 e (cost of carbon abatement, or cost of carbon reduction) using data for the year 2021. This study is focused on a heavy-duty off-road material handler, the Pettibone Cary-Lift 204i. This machine’s specialty application, including events like unloading large oil pipes from a railcar, requires a unique electrified architecture that suits its specific needs. However, the results from this study may be extrapolated to similar machinery to inform fuel savings options across the heavy-duty off-road industry. In this study, a unique electrified architecture is determined for the Cary-Lift. This architecture is informed by multiple rounds of a Pugh matrix decision analysis to select a shortened list of desirable electrified architectures. The shortened list is modeled and simulated to determine CO 2 e, cost, and cost per CO 2 e. A final architecture is determined as a plug-in series hybrid that reduces fuel consumption by 65%, targeting the large fuel and CO 2 e savings that are likely to be required for the future of the heavy-duty off-road industry.

33 ADVANCED PROPULSION SYSTEMS↗

Pele: An Exascale-Ready Suite of Combustion Codes

High fidelity simulations of realistic combustion devices are extremely demanding computationally because of the requirements to capture complex fuel chemical decomposition, its intricate interactions with turbulent, often multiphase, flows, and the wide separation of space and time scales between the thin flame and the device boundaries. Software required to carry out such computations tends to be extremely complex, particularly when designed to exploit hardware accelerators, and can be difficult to port and maintain. We present Pele, a performance portable suite of tools for the simulation of combustion systems, including codes to evolve reactive multiphase configurations in the low Mach number and compressible flow regimes, along with a set of inter-compatible post processing and in situ analysis tools. The Pele suite of tools is built on top of the AMReX framework for block-structured adaptive mesh refinement, which provides efficient data structures and algorithms that enable the development of a wide variety of efficient mesh and particle based PDE integration schemes. A hierarchical MPI+X parallelism scheme supports CPU-only and accelerated architectures, where X can be OpenMP, CUDA, and HIP based approaches for intra-node computational work distribution. The algorithms and data structures underlying the Pele simulation and analysis tools are highly scalable and performant across a wide variety of high-performance computing platforms, including DOEs newest exascale-class machines, Frontier and Aurora. The simulation and analysis tools are fully documented and freely distributed as open source via GitHub. We present key algorithmic and software challenges, solution strategies, performance and resulting set of capabilities.

AMReX↗

Integrating ytopt and libEnsemble to autotune OpenMC

Ytopt is a Python machine-learning-based autotuning software package developed within the ECP PROTEAS-TUNE project. The ytopt software adopts an asynchronous search framework that consists of sampling a small number of input parameter configurations and progressively fitting a surrogate model over the input-output space until exhausting the user-defined maximum number of evaluations or the wall-clock time. libEnsemble is a Python toolkit for coordinating workflows of asynchronous and dynamic ensembles of calculations across massively parallel resources developed within the ECP PETSc/TAO project. libEnsemble helps users take advantage of massively parallel resources to solve design, decision, and inference problems and expands the class of problems that can benefit from increased parallelism. In this paper we present our methodology and framework to integrate ytopt and libEnsemble to take advantage of massively parallel resources to accelerate the autotuning process. Specifically, we focus on using the proposed framework to autotune the ECP ExaSMR application OpenMC, an open source Monte Carlo particle transport code. OpenMC has seven tunable parameters some of which have large ranges such as the number of particles in-flight, which is in the range of 100,000 to 8 million, with its default setting of 1 million. Setting the proper combination of these parameter values to achieve the best performance is extremely time-consuming. Therefore, we apply the proposed framework to autotune the MPI/OpenMP offload version of OpenMC based on a user-defined metric such as the figure of merit (FoM) (particles/s) or energy efficiency energy-delay product (EDP) on Crusher at Oak Ridge Leadership Computing Facility. In conclusion, the experimental results show that we achieve the improvement up to 29.49% in FoM and up to 30.44% in EDP.

Autotuning↗

Quantum model learning agent: characterisation of quantum systems through machine learning

Accurate models of real quantum systems are important for investigating their behaviour, yet are difficult to distil empirically. Here, we report an algorithm—the quantum model learning agent (QMLA)—to reverse engineer Hamiltonian descriptions of a target system. We test the performance of QMLA on a number of simulated experiments, demonstrating several mechanisms for the design of candidate Hamiltonian models and simultaneously entertaining numerous hypotheses about the nature of the physical interactions governing the system under study. QMLA is shown to identify the true model in the majority of instances, when provided with limited a priori information, and control of the experimental setup. Our protocol can explore Ising, Heisenberg and Hubbard families of models in parallel, reliably identifying the family which best describes the system dynamics. We demonstrate QMLA operating on large model spaces by incorporating a genetic algorithm to formulate new hypothetical models. The selection of models whose features propagate to the next generation is based upon an objective function inspired by the Elo rating scheme, typically used to rate competitors in games such as chess and football. In all instances, our protocol finds models that exhibit F 1 score ≥ 0.88 when compared with the true model, and it precisely identifies the true model in 72% of cases, whilst exploring a space of over 250 000 potential models. By testing which interactions actually occur in the target system, QMLA is a viable tool for both the exploration of fundamental physics and the characterisation and calibration of quantum devices.

97 MATHEMATICS AND COMPUTING↗

Optimizing temperature distributions for training neural quantum states using parallel tempering

Parametrized artificial neural networks (ANNs) can be very expressive ansatzes for variational algorithms, reaching state-of-the-art energies on many quantum many-body Hamiltonians. Nevertheless, the training of the ANN can be slow and stymied by the presence of local minima in the parameter landscape. One approach to mitigate this issue is to use parallel tempering methods, and in this work, we focus on the role played by the temperature distribution of the parallel tempering replicas. Using an adaptive method that adjusts the temperatures in order to equate the exchange probability between neighboring replicas, we show that this temperature optimization can significantly increase the success rate of the variational algorithm with negligible computational cost by eliminating bottlenecks in the replicas' random walk. Furthermore, we demonstrate this using two different neural networks, a restricted Boltzmann machine and a feedforward network, which we use to study a toy problem based on a permutation invariant Hamiltonian with a pernicious local minimum and the 𝐽 1 −𝐽 2 model on a rectangular lattice.

Neural network simulations↗

Performance Analysis of an Optimization Algorithm for Metamaterial Design on the Integrated High-Performance Computing and Quantum Systems

Optimizing metamaterials with complex geometries is a big challenge. Although an active learning algorithm, combining machine learning (ML), quantum computing, and optical simulation, has emerged as an efficient optimization tool, it still faces difficulties in optimizing complex structures that have potentially high performance. In this work, we comprehensively analyze the performance of an optimization algorithm for metamaterial design on the integrated HPC and quantum systems. We demonstrate significant time advantages through message-passing interface (MPI) parallelization on the high-performance computing (HPC) system showing approximately 54% faster ML tasks and 67 times faster optical simulation against serial workloads. Furthermore, we analyze the performance of a quantum algorithm designed for optimization, which runs with various quantum simulators on a local computer or HPC-quantum system. Results showcase ~24 times speedup when executing the optimization algorithm on the HPC-quantum hybrid system. This study paves a way to optimize complex metamaterials using the integrated HPC-quantum system.

Kim, Seongmin↗

Asynchronous-many-task systems: Challenges and opportunities - Scaling an AMR astrophysics code on exascale machines using Kokkos and HPX

Dynamic and adaptive mesh refinement is pivotal in high-resolution, multi-physics, multi-model simulations, necessitating precise physics resolution in localized areas across expansive domains. Today’s supercomputers’ extreme heterogeneity presents a significant challenge for dynamically adaptive codes, highlighting the importance of achieving performance portability at scale. Our research focuses on astrophysical simulations, particularly stellar mergers, to elucidate early universe dynamics. Here, we present Octo-Tiger, leveraging Kokkos, HPX, and SIMD for portable performance at scale in complex, massively parallel adaptive multi-physics simulations. Octo-Tiger supports diverse processors, accelerators, and network backends. Experiments demonstrate exceptional scalability across several heterogeneous supercomputers including Perlmutter, Frontier, and Fugaku, encompassing major GPU architectures and x86, ARM, and RISC-V CPUs. Parallel efficiency of 47.59% (110,080 cores and 6880 hybrid A100 GPUs) on a full-system run on Perlmutter (26% HPCG peak performance) and 51.37% (using 32,768 cores and 2048 MI250X) on Frontier are achieved.

97 MATHEMATICS AND COMPUTING↗

AENET–LAMMPS and AENET–TINKER : Interfaces for accurate and efficient molecular dynamics simulations with machine learning potentials

Machine-learning potentials (MLPs) trained on data from quantum-mechanics based first-principles methods can approach the accuracy of the reference method at a fraction of the computational cost. To facilitate efficient MLP-based molecular dynamics and Monte Carlo simulations, an integration of the MLPs with sampling software is needed. Here, we develop two interfaces that link the atomic energy network (ænet) MLP package with the popular sampling packages TINKER and LAMMPS. The three packages, ænet, TINKER, and LAMMPS, are free and open-source software that enable, in combination, accurate simulations of large and complex systems with low computational cost that scales linearly with the number of atoms. Scaling tests show that the parallel efficiency of the ænet–TINKER interface is nearly optimal but is limited to shared-memory systems. The ænet–LAMMPS interface achieves excellent parallel efficiency on highly parallel distributed memory systems and benefits from the highly optimized neighbor list implemented in LAMMPS. We demonstrate the utility of the two MLP interfaces for two relevant example applications: the investigation of diffusion phenomena in liquid water and the equilibration of nanostructured amorphous battery materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

FuseIM: Fusing Probabilistic Traversals for Influence Maximization on Exascale Systems

Probabilistic breadth-first traversals (BPTs) are used in many network science and graph machine learning applications. In this paper, we are motivated by the application of BPTs in stochastic diffusion-based graph problems such as influence maximization. These applications heavily rely on BPTs to implement a Monte-Carlo sampling step for their approximations. Given the large sampling complexity, stochasticity of the diffusion process, and the inherent irregularity in real-world graph topologies, efficiently parallelizing these BPTs remains significantly challenging. In this paper, we present a new algorithm to fuse massive number of concurrently executing BPTs with random starts on the input graph. Our algorithm is designed to fuse BPTs by combining separate traversals into a unified frontier on distributed multi-GPU systems. To show the general applicability of the fused BPT technique, we have incorporated it into two state-of-the-art influence maximization parallel implementations (gIM and Ripples). Our experiments on up to 4K nodes of the OLCF Frontier supercomputer (32,768 GPUs and 196K CPU cores) show strong scaling behavior, and that fused BPTs can improve the performance of these implementations up to 34x (for gIM) and ~360x (for Ripples).

Neff, Reece W.↗

Data-Driven Optimization of the Processing Window for 316H Components Fabricated Using Laser Powder Bed Fusion

The Advanced Materials and Manufacturing Technologies Program is focused on accelerating the development and deployment of advanced materials and components fabricated via additive manufacturing with a specific focus on laser powder bed fusion (LPBF). As an initial case study, the program has selected 316H stainless steel (SS) as an initial material around which to develop a code case development strategy. This strategy involves two parallel approaches: (1) an equivalency approach whereby round-robin testing across multiple collaborating laboratories demonstrates repeatability in processing and direct comparisons with conventional wrought 316H material and (2) a revolutionary approach to code qualification combining in situ data collection and high-fidelity modeling to capture, predict, and bound the performance of LPBF 316HSS components. As part of this campaign, this work package has initiated an extensive process optimization campaign across three laboratories, each printing variations of LPBF 316HSS using three different LPBF units (Concept Laser, EOS, and Renishaw). In FY23, ORNL has focused on unique experimental designs spanning wide ranges in energy inputs and turning knobs such as scan speed, laser power, hatch spacing, layer thickness, spot size, scan rotation, and more. On the Concept Laser M2, 72 different combinations of processing variables were investigated with duplicate samples and different powder compositions. In total, 252 samples were printed with combined in situ sensing data. A parallel design of experiments was conducted on the Renishaw AM400 with an additional 390 printed specimens for analysis. All 642 miniature specimens, each with unique features included in each print to capture geometry-related heterogeneity, were subjected to high-throughput x-ray computed tomography (XCT) analysis to enable the downselection of specific processing parameters of interest. Then, using electrical discharge machining (EDM), miniature tensile specimens were extracted for mechanical testing and microscopy investigations. From the analysis performed in FY23, it was found that powder composition drastically affects the resulting microstructure and mechanical performance of 316SS. Specifically, changing from 316L to 316HSS powder results in a wide range of grain sizes with varying degrees of preferred grain orientation, which increases as a function of energy density. It was also found that due to stored heat in thin fin–type features, large microstructural differences can be seen within one part printed with one set of processing parameters. These variations in microstructure features, including grain size, the nanoscale dislocation structure, and grain texture, will all affect the irradiation performance and high-temperature mechanical performance of LPBF 316HSS parts. Two sets of concept laser processing parameters, spanning both refined and columnar grain structures, were scaled to print larger 316H builds for campaign testing (high-temperature creep and irradiation). In addition, at least two optimized processing parameter sets were identified for the Renishaw AM400 for round-robin testing in FY24 with Argonne National Laboratory. Future work includes printing samples using identical parameters identified by partner institutions, providing material for corrosion and high-temperature mechanical testing, and continuing evaluations of heterogeneity in larger printed parts.

36 MATERIALS SCIENCE↗

Data-Driven Optimization of the Processing Window for 316H Components Fabricated Using Laser Powder Bed Fusion

The Advanced Materials and Manufacturing Technologies Program is focused on accelerating the development and deployment of advanced materials and components fabricated via additive manufacturing with a specific focus on laser powder bed fusion (LPBF). As an initial case study, the program has selected 316H stainless steel (SS) as an initial material around which to develop a code case development strategy. This strategy involves two parallel approaches: (1) an equivalency approach whereby round-robin testing across multiple collaborating laboratories demonstrates repeatability in processing and direct comparisons with conventional wrought 316H material and (2) a revolutionary approach to code qualification combining in situ data collection and high-fidelity modeling to capture, predict, and bound the performance of LPBF 316HSS components. As part of this campaign, this work package has initiated an extensive process optimization campaign across three laboratories, each printing variations of LPBF 316HSS using three different LPBF units (Concept Laser, EOS, and Renishaw). In FY23, ORNL has focused on unique experimental designs spanning wide ranges in energy inputs and turning knobs such as scan speed, laser power, hatch spacing, layer thickness, spot size, scan rotation, and more. On the Concept Laser M2, 72 different combinations of processing variables were investigated with duplicate samples and different powder compositions. In total, 252 samples were printed with combined in situ sensing data. A parallel design of experiments was conducted on the Renishaw AM400 with an additional 390 printed specimens for analysis. All 642 miniature specimens, each with unique features included in each print to capture geometry-related heterogeneity, were subjected to high-throughput x-ray computed tomography (XCT) analysis to enable the downselection of specific processing parameters of interest. Then, using electrical discharge machining (EDM), miniature tensile specimens were extracted for mechanical testing and microscopy investigations. From the analysis performed in FY23, it was found that powder composition drastically affects the resulting microstructure and mechanical performance of 316SS. Specifically, changing from 316L to 316HSS powder results in a wide range of grain sizes with varying degrees of preferred grain orientation, which increases as a function of energy density. It was also found that due to stored heat in thin fin–type features, large microstructural differences can be seen within one part printed with one set of processing parameters. These variations in microstructure features, including grain size, the nanoscale dislocation structure, and grain texture, will all affect the irradiation performance and high-temperature mechanical performance of LPBF 316HSS parts. Two sets of concept laser processing parameters, spanning both refined and columnar grain structures, were scaled to print larger 316H builds for campaign testing (high-temperature creep and irradiation). In addition, at least two optimized processing parameter sets were identified for the Renishaw AM400 for round-robin testing in FY24 with Argonne National Laboratory. Future work includes printing samples using identical parameters identified by partner institutions, providing material for corrosion and high-temperature mechanical testing, and continuing evaluations of heterogeneity in larger printed parts.

36 MATERIALS SCIENCE↗

Impurity transport in PISCES-RF

Linear plasma devices (LPD) utilizing a helicon plasma source, a high density light ion source, can generate impurities due to progressive erosion of the radio frequency (RF) transmission window caused by rectified sheath voltage. These source-born impurities can entrain and be transported by the plasma toward a target, affecting plasma-material interaction studies. Earlier work on material testing in Prototype-Materials Plasma Exposure eXperiment at ORNL revealed significant source impurity deposition on downstream targets. However, using a similar RF source, no target impurity deposition is observed in Plasma Interaction Surface Component Experimental Station (PISCES)-RF despite evidence of RF window erosion in the source region, thereby motivating the present work. Experimentally, using various magnetic field configurations upstream of the PISCES-RF plasma source and seeding titanium (Ti) impurities at various axial locations, impurity transport and deposition along the machine axis were investigated. It was found that Ti deposition was localized to the side of the plasma source where the Ti impurity was seeded. In contrast, aluminum (Al) deposition, originating from the sputtering of the helicon window, occurred predominantly upstream of the plasma source, suggesting an asymmetry in the axial transport of eroded RF window material. These observations suggest a stagnation of the parallel plasma flow immediately downstream of the plasma source, with impurity ions remaining unmagnetized near the source upstream. Al deposition in magnetic field-free regions in PISCES-RF indicates that sputtered Al impurities likely remained neutral due to their large ionization mean-free path under PISCES-RF conditions. Plasma modeling and simulation supported this, indicating that Al-neutrals transport toward the helicon source upstream for low electron density cases. It was found that the Larmor radius of the Al ions was greater than the plasma radius towards the source upstream and remained weakly magnetized in PISCES-RF, meaning that plasma source-born impurities are not efficiently entrained in the plasma flow. These findings provide critical insights into impurity transport in helicon plasma-based LPDs.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Parallel physics-informed neural networks via domain decomposition

Here we develop a distributed framework for the physics-informed neural networks (PINNs) based on two recent extensions, namely conservative PINNs (cPINNs) and extended PINNs (XPINNs), which employ domain decomposition in space and in time-space, respectively. This domain decomposition endows cPINNs and XPINNs with several advantages over the vanilla PINNs, such as parallelization capacity, large representation capacity, efficient hyperparameter tuning, and is particularly effective for multi-scale and multi-physics problems. Here, we present a parallel algorithm for cPINNs and XPINNs constructed with a hybrid programming model described by MPI + X, where X ∈ {CPUs, GPUs}. The main advantage of cPINN and XPINN over the more classical data and model parallel approaches is the flexibility of optimizing all hyperparameters of each neural network separately in each subdomain. We compare the performance of distributed cPINNs and XPINNs for various forward problems, using both weak and strong scalings. Our results indicate that for space domain decomposition, cPINNs are more efficient in terms of communication cost but XPINNs provide greater flexibility as they can also handle time-domain decomposition for any differential equations, and can deal with any arbitrarily shaped complex subdomains. To this end, we also present an application of the parallel XPINN method for solving an inverse diffusion problem with variable conductivity on the United States map, using ten regions as subdomains.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

DIVA/DeviceEditor (DIVA) v6.0.0

The DIVA software interfaces a process in which researchers design their DNA with a web-based graphical user interface (DeviceEditor), submit their designs to a central queue, and a few weeks later receive their sequence-verified clonal constructs. Each researcher independently designs the DNA to be constructed with a web-based BioCAD tool, and presses a button to submit their designs to a central queue. Researchers have web-based access to their DNA design queues, and can track the progress of their submitted designs as they progress from "evaluation", to "waiting for reagents", to "in progress", to "complete". Researchers access their completed constructs through the central DNA repository. Along the way, all DNA construction success/failure rates are captured in a central database. Once a design has been submitted to the queue, a small number of dedicated staff evaluate the design for feasibility and provide feedback to the responsible researcher if the design is either unreasonable (e.g., encompasses a combinatorial library of a billion constructs) or small design changes could significantly facilitate the downstream implementation process. The dedicated staff then use DNA assembly design automation software to optimize the DNA construction process for the design, leveraging existing parts from the DNA repository where possible and ordering synthetic DNA where necessary. Once all requisite process inputs are available, the design progresses from "waiting for reagents" to "in progress" in the design queue. Human-readable and machine-parseable DNA construction protocols output by the DNA assembly design automation software are then executed by the dedicated staff exploiting lab automation devices wherever possible. Since the all employed DNA construction methods are sequence-agnostic, standardized (utilize the same enzymatic master mixes and reaction conditions), completely independent DNA construction tasks can be aggregated into the same multi-well plates and pursued in parallel. The resulting sets of cloned constructs can then be screened by high-throughput next-gen sequencing platforms for sequence correctness. A combination of long read-length (e.g., PacBio) and paired-end read platforms (e.g., Illumina) would be exploited depending the particular task at hand (e.g., PacBio might be sufficient to screen a set of pooled constructs with significant gene divergence). Post sequence verification, designs for which at least one correct clone was identified will progress to a "complete" status, while designs for which no correct clones were identified will progress to a "failure" status. Depending on the failure mode (e.g., no transformants), and how many prior attempts/variations of assembly protocol have been already made for a given design, subsequent attempts may be made or the design can progress to a "permanent failure" state. All success and failure rate information will be captured during the process, including at which stage a given clonal construction procedure failed (e.g., no PCR product) and what the exact failure was (e.g. assembly piece 2 missing). This success/failure rate data can be leveraged to refine the DNA assembly design process.

Plahar, Hector↗

MAPredict: Static Analysis Driven Memory Access Prediction Framework for Modern CPUs

Application memory access patterns are crucial in deciding how much traffic is served by the cache and forwarded to the dynamic random-access memory (DRAM). However, predicting such memory traffic is difficult because of the interplay of prefetchers, compilers, parallel execution, and innovations in manufacturer-specific micro-architectures. This research introduced MAPredict, a static analysis-driven framework that addresses these challenges to predict last-level cache (LLC)-DRAM traffic. By exploring and analyzing the behavior of modern Intel processors, MAPredict formulates cache-aware analytical models. MAPredict invokes these models to predict LLC-DRAM traffic by combining the application model, machine model, and user-provided hints to capture dynamic information. MAPredict successfully predicts LLC-DRAM traffic for different regular access patterns and provides the means to combine static and empirical observations for irregular access patterns. Evaluating 130 workloads from six applications on recent Intel micro-architectures, MAPredict yielded an average accuracy of 99% for streaming, 91% for strided, and 92% for stencil patterns. By coupling static and empirical methods, up to 97% average accuracy was obtained for random access patterns on different micro-architectures.

Monil, M. A. H.↗

Silicon carbide power inverter/rectifier for electric machines

The present disclosure involves a two stage inverter, a system for electrical power conversation, and a method of converting electrical power using silicon carbide (SiC) metal-oxide-semiconductor field-effect transistors (MOSFETs). One example implementation includes using two or more SiC MOSFETs in series with each MOSFET having a gate terminal for triggering a state switch between an on (conducting) and off (non-conducting) state of the MOSFET. An AC terminal is connected between the series SiC MOSFETS, and the series SiC MOSFETs are connected across a DC bus and in parallel with one or more capacitors.

Shenoy, Suratkal P.↗

I/O in Machine Learning Applications on HPC Systems: A 360-degree Survey

Growing interest in Artificial Intelligence (AI) has resulted in a surge in demand for faster methods of Machine Learning (ML) model training and inference. This demand for speed has prompted the use of high performance computing (HPC) systems that excel in managing distributed workloads. Because data is the main fuel for AI applications, the performance of the storage and I/O subsystem of HPC systems is critical. In the past, HPC applications accessed large portions of data written by simulations or experiments or ingested data for visualizations or analysis tasks. ML workloads perform small reads spread across a large number of random files. This shift of I/O access patterns poses several challenges to modern parallel storage systems. In this paper, we survey I/O in ML applications on HPC systems, and target literature within a 6-year time window from 2019 to 2024. We define the scope of the survey, provide an overview of the common phases of ML, review available profilers and benchmarks, examine the I/O patterns encountered during offline data preparation, training, and inference, and explore I/O optimizations utilized in modern ML frameworks and proposed in recent literature. Lastly, we seek to expose research gaps that could spawn further R&D.

97 MATHEMATICS AND COMPUTING↗

Solving the electronic structure problem for over 100000 atoms in real space

Using a real-space high-order finite-difference approach, we investigate the electronic structure of large spherical silicon nanoclusters. Within Kohn-Sham density functional theory and using pseudopotentials, we report the self-consistent field convergence of a system with over 100000 atoms: a Si 107,641 ⁢H 9,084 nanocluster with a diameter of 16 nm. Our approach uses Chebyshev-filtered subspace iteration to speed up the convergence of the eigenspace, and blockwise Hilbert space-filling curves to speed up sparse matrix-vector multiplications, all of which are implemented in the parsec code. For the largest system, we utilized 2048 nodes (114 688 cores) on the Frontera machine in the Texas Advanced Computing Center. Our quantitative analysis of the electronic structure shows how it gradually approaches its bulk counterpart as a function of nanocluster size. The band gap is enlarged due to quantum confinement in nanoclusters, but decreases as the system size increases, as expected. In conclusion, our work serves as a proof of concept for the capacity of the real-space approach in efficiently parallelizing very large calculations using high-performance computer platforms, which can straightforwardly be replicated in other systems with more than 10 5 atoms.

0-dimensional systems↗