Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Scalable quantum computer”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

359 records · Page 20

Qubit Assignment Using Time Reversal

As quantum computers with large numbers of qubits become increasingly available, experiments executed on a given device may not utilize all available qubits. In this case, the outcome of executing a quantum program will depend on the ability to efficiently select a subset of high-performing physical qubits. For any given quantum program and device there are many ways to assign physical qubits for execution of the program, and assignments will differ in performance due to the variability in quality across qubits and entangling operations on a single device. Evaluating the performance of each assignment using fidelity estimation introduces significant experimental overhead and will be infeasible for many applications, while relying on standard device benchmarks provides incomplete information about the performance of any specific program. Furthermore, the number of possible assignments grows combinatorially in the number of qubits on the device and in the program, motivating the use of heuristic optimization techniques. We demonstrate a practical solution to the problem of qubit assignment by using simulated annealing with a cost function based on the Loschmidt echo, a diagnostic that measures the reversibility of a quantum process. We provide theoretical justification for this choice of cost function by demonstrating that the optimal qubit assignment coincides with the optimal qubit assignment based on state fidelity in the weak error limit, and we provide experimental justification using diagnostics performed on Google’s superconducting qubit devices. We then establish the performance of simulated annealing for qubit assignment using classical simulations of noisy devices as well as optimization experiments performed on a quantum processor. Our results demonstrate that the use of Loschmidt echoes and simulated annealing provides a scalable and flexible approach to optimizing qubit assignment on near-term hardware.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Universal Monte Carlo Event Generator

With the Jefferson Lab 12 GeV physics program underway and plans for the future Electron-Ion Collider (EIC), the nuclear physics community is entering a new era of exploration of QCD phenomena involving extensive data taking and event-level processing. This brings with it the potential for unprecedented access to multidimensional particle momentum distributions (PMDs) that can be connected to various theoretical frameworks by unfolding the emergent quantum mechanical properties of QCD using the PMDs. In practice, the PMDs are rendered as discretized histograms (typically one- or two-dimensional projections), and detector effects must be taken into account to unfold the pure detector effect-free PMDs that can be connected with theory. One of the challenges in this new era is obtaining faithful reconstructions of the multidimensional PMDs that preserve all of the inherent particle correlations. In this LDRD project we developed a novel approach using machine learning (ML) that solves this challenge, by avoiding entirely the need to use histograms as the main numerical technique to obtain the detector effect-free PMDs. The new approach is, moreover, scalable to higher dimensional PMDs. The central idea involves training neural networks (NNs) to generate synthetic event-level data (momenta 4-vectors of final state particles) to preserve all correlations among the particles. This is achieved by converting the trial synthetic vertex-level events to detector-level events using detector simulators. The NNs are then tuned using a specialized distance metric between the synthetic detector events and the real detector events.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Enhancing scalability and accuracy of quantum poisson solver

The Poisson equation has many applications across the broad areas of science and engineering. Most quantum algorithms for the Poisson solver presented so far either suffer from lack of accuracy and/or are limited to very small sizes of the problem and thus have no practical usage. In this regard, our previous work showed a proof-of-concept demonstration in advancing quantum Poisson solver algorithm and validated preliminary results for a simple case of 3 x 3 problem. In this work, we delve into comprehensive research details, presenting the results on up to 15 x 15 problems that include step-by-step improvements in Poisson equation solutions, scaling performance, and experimental exploration. In particular, we demonstrate the implementation of eigenvalue amplification by a factor of up to 2 8 , achieving a significant improvement in the accuracy of our quantum Poisson solver and comparing that to the exact solution. Additionally, we present success probability results, highlighting the reliability of our quantum Poisson solver. Moreover, we explore the scaling performance of our algorithm against the circuit depth and width, demonstrating how our approach scales with larger problem sizes and thus further solidifies the practicality of easy adaptation of this algorithm in real-world applications. We also discuss a multilevel strategy for how this algorithm might be further improved to explore much larger problems with greater performance. Finally, through our experiments on the IBM quantum hardware, we conclude that though overall results on the existing NISQ hardware are dominated by the error in the CNOT gates, this work opens a path to realizing a multidimensional Poisson solver on near-term quantum hardware.

97 MATHEMATICS AND COMPUTING↗

Coarse-Grained Density Functional Theory Predictions via Deep Kernel Learning

Scalable electronic predictions are critical for soft materials design. Recently, the Electronic Coarse-Graining (ECG) method was introduced to renormalize all-atom quantum chemical (QC) predictions to coarse-grained (CG) resolutions using deep neural networks (DNNs). While DNNs can learn complex representations that prove challenging for kernel-based methods, they are susceptible to overfitting and the overconfidence of uncertainty estimations. Here, we develop ECG within a GPU-accelerated Deep Kernel Learning (DKL) framework to enable CG QC predictions using range-separated hybrid density functional theory (DFT), obtaining a 107 speedup relative to naive all-atom QC. By treating the predicted electronic properties as random Gaussian Processes, DKL incorporates CG mapping degeneracy by learning the distribution of electronic energies as a function of CG configuration. DKL-ECG accurately reproduces molecular orbital energies from range-separated DFT while facilitating efficient training via active learning using the uncertainties provided by DKL. Further, we show that while active learning algorithms enable efficient sampling of a more diverse configurational space relative to random sampling, all explored query methods exhibit comparable performance for the examined system. We attribute this result to the significant overlap of the feature space and output property distributions across multiple temperatures.

97 MATHEMATICS AND COMPUTING↗

Runtime performance of a GAMESS quantum chemistry application offloaded to GPUs

Summary Computational chemistry is at the forefront of solving urgent societal problems, such as polymer upcycling and carbon capture. The complexity of modeling these processes at appropriate length and time scales is mainly manifested in the number and types of chemical species involved in the reactions and may require models of several thousand atoms and large basis sets to accurately capture the chemical complexity and heterogeneity in the physical and chemical processes. The quantum chemistry package General Atomic and Molecular Electronic Structure System (GAMESS) has a wide array of methods that can efficiently and accurately treat complex chemical systems. In this work, we have used the GAMESS Effective Fragment Molecule Orbital (EFMO) method for electronic structure calculation of a challenging mesoporous silica nanoparticle (MSN) model surrounded by about 4700 water molecules to investigate the strong scaling and GPU offloading on hybrid CPU‐GPU nodes. Experiments were performed on the Perlmutter platform at the National Energy Research Scientific Computing Center. Good strong scaling and load balancing have been observed on up to 88 hybrid nodes for different settings of the execution parameters for the calculation considered here. When GPUs are oversubscribed by offloading work from multiple CPU processes, using the NVIDIA multi‐process service (MPS) has consistently reduced time to solution and energy consumed. Additionally, for some configuration parameter settings, oversubscription with MPS improved performance by up to 5.8% over the case without oversubscription.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

QPatLib v1.0 — Measurement-based quantum simulation Pauli string unitary pattern collections

This Zenodo record accompanies the paper “Scalable Measurement-Based Quantum Simulation Patterns for Benchmarking” arXiv.2605.12502 and provides QPatLib v1.0 measurement-pattern datasets in human-readable JSONL together with a ZIP archive of OpenQASM 3.0 circuits used for validation and reproducibility. The patterns and circuits implement Pauli string unitaries for benchmark cases. Cases include all possible string combinations for less than 6 qubits and strings used in Hamiltonians for certain diatomic molecules for 6 or more qubits. Format: Each pattern_*.jsonl file is containins measurement patterns for all subsets for a given model/instance and subset strategy: it begins with a preamble containing model metadata, subset definitions, provenance, and (when feasible) full-pattern test results, followed by one pattern entry per subset. Each subset entry includes a required pattern_ascii field storing the measurement pattern in the measurement-calculus/Graphix standard with signal shifting, written left-to-right in the canonical order nodes → edges → measurements (with signal dependencies) → byproduct corrections (X/Z). The circuits are included as circuit_files.zip. Patterns in this record were validated against the corresponding circuits and checked for causal flow. Codes for generating these patterns can be found at QPatLib repository on Github

Graphix↗

Diagnosing Barren Plateaus with Tools from Quantum Optimal Control

Variational Quantum Algorithms (VQAs) have received considerable attention due to their potential for achieving near-term quantum advantage. However, more work is needed to understand their scalability. One known scaling result for VQAs is barren plateaus, where certain circumstances lead to exponentially vanishing gradients. It is common folklore that problem-inspired ansatzes avoid barren plateaus, but in fact, very little is known about their gradient scaling. In this work we employ tools from quantum optimal control to develop a framework that can diagnose the presence or absence of barren plateaus for problem-inspired ansatzes. Such ansatzes include the Quantum Alternating Operator Ansatz (QAOA), the Hamiltonian Variational Ansatz (HVA), and others. With our framework, we prove that avoiding barren plateaus for these ansatzes is not always guaranteed. Specifically, we show that the gradient scaling of the VQA depends on the degree of controllability of the system, and hence can be diagnosed through the dynamical Lie algebra $\mathfrak{g}$ obtained from the generators of the ansatz. We analyze the existence of barren plateaus in QAOA and HVA ansatzes, and we highlight the role of the input state, as different initial states can lead to the presence or absence of barren plateaus. Taken together, our results provide a framework for trainability-aware ansatz design strategies that do not come at the cost of extra quantum resources. Moreover, we prove no-go results for obtaining ground states with variational ansatzes for controllable system such as spin glasses. Our work establishes a link between the existence of barren plateaus and the scaling of the dimension of $\mathfrak{g}$.

97 MATHEMATICS AND COMPUTING↗

Implementation of scalable suspended superinductors

Superinductors have become a crucial component in the superconducting circuit toolbox, playing a key role in the development of more robust qubits. Enhancing the performance of these devices can be achieved by suspending the superinductors from the substrate, thereby reducing stray capacitance. Here, we present a fabrication framework for constructing superconducting circuits with suspended superinductors in planar architectures. To validate the effectiveness of this process, we systematically characterize both resonators and qubits with suspended arrays of Josephson junctions, ultimately confirming the high quality of the superinductive elements. In addition, this process is broadly compatible with other types of superinductors and circuit designs. Furthermore, our results not only pave the way for scalable superconducting architectures utilizing superinductors but also provide the primitive for future investigation of loss mechanisms associated with the device substrate.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Silicon nitride stress-optic microresonator modulator for optical control applications

Modulation-based control and locking of lasers, filters and other photonic components is a ubiquitous function across many applications that span the visible to infrared (IR), including atomic, molecular and optical (AMO), quantum sciences, fiber communications, metrology, and microwave photonics. Today, modulators used to realize these control functions consist of high-power bulk-optic components for tuning, sideband modulation, and phase and frequency shifting, while providing low optical insertion loss and operation from DC to 10s of MHz. In order to reduce the size, weight and cost of these applications and improve their scalability and reliability, modulation control functions need to be implemented in a low loss, wafer-scale CMOS-compatible photonic integration platform. The silicon nitride integration platform has been successful at realizing extremely low waveguide losses across the visible to infrared and components including high performance lasers, filters, resonators, stabilization cavities, and optical frequency combs. Yet, progress towards implementing low loss, low power modulators in the silicon nitride platform, while maintaining wafer-scale process compatibility has been limited. Here we report a significant advance in integration of a piezo-electric (PZT, lead zirconate titanate) actuated micro-ring modulation in a fully-planar, wafer-scale silicon nitride platform, that maintains low optical loss (0.03 dB/cm in a 625 µm resonator) at 1550 nm, with an order of magnitude increase in bandwidth (DC - 15 MHz 3-dB and DC - 25 MHz 6-dB) and order of magnitude lower power consumption of 20 nW improvement over prior PZT modulators. The modulator provides a >14 dB extinction ratio (ER) and 7.1 million quality-factor (Q) over the entire 4 GHz tuning range, a tuning efficiency of 162 MHz/V, and delivers the linearity required for control applications with 65.1 dB·Hz 2/3 and 73.8 dB·Hz 2/3 third-order intermodulation distortion (IMD3) spurious free dynamic range (SFDR) at 1 MHz and 10 MHz respectively. We demonstrate two control applications, laser stabilization in a Pound-Drever Hall (PDH) lock loop, reducing laser frequency noise by 40 dB, and as a laser carrier tracking filter. This PZT modulator design can be extended to the visible in the ultra-low loss silicon nitride platform with minor waveguide design changes. This integration of PZT modulation in the ultra-low loss silicon nitride waveguide platform enables modulator control functions in a wide range of visible to IR applications such as atomic and molecular transition locking for cooling, trapping and probing, controllable optical frequency combs, low-power external cavity tunable lasers, quantum computers, sensors and communications, atomic clocks, and tunable ultra-low linewidth lasers and ultra-low phase noise microwave synthesizers.

Wang, Jiawei (ORCID:0000000257965220)↗

NWChem: Past, present, and future

Specialized computational chemistry packages have permanently reshaped the landscape of chemical sciences by providing tools to support and guide the experimental effort and for prediction of chemical and materials properties. In this regard, a special role has been played by electronic structure packages where complex chemical and materials processes can be modeled using first-principle-driven methodologies. Over the last few decades, the rapid development of computing technologies and a tremendous increase in computational power has offered a unique chance to study complex chemical transformations using sophisticated and predictive many-body techniques to describe correlated behavior of electrons in molecular and condensed phase systems at different levels of theory. In enabling these simulations, a critical role has been played by novel parallel algorithms capable of taking advantage of computational resources to address polynomial scaling of electronic structure methods. NWChem was among the first electronic structure codes that focused on delivering scalable performance for electronic structure simulations. Herein, we briefly review the NWChem suite of computational codes including its history, design principles, parallel tools, current capabilities, outreach and outlook.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Scalable Experimental Bounds for Entangled Quantum State Fidelities

Estimating the state preparation fidelity of highly entangled states on noisy intermediate-scale quantum (NISQ) devices is important for benchmarking and application considerations. Unfortunately, exact fidelity measurements quickly become prohibitively expensive, as they scale exponentially as O(3 N for N-qubit states, using full state tomography with measurements in all Pauli bases combinations. However, Somma et al.established that the complexity could be drastically reduced when looking at fidelity lower bounds for states that exhibit symmetries, such as Dicke states and GHZ states. These bounds must still be tight enough for larger states to provide reasonable estimations on NISQ devices. For the first time and more than 15 years after the theoretical introduction, we report meaningful lower bounds for the state preparation fidelity of all Dicke states up to N=10 and all GHZ states up to N=20 on Quantinuum H1 ion-trap systems using efficient implementations of recently proposed scalable circuits for these states. Our achieved lower bounds match or exceed previously reported exact fidelities on superconducting systems for much smaller states. Furthermore, we provide evidence that for large Dicke states |$D^{N}_{N/2}\rangle$, we may resort to a GHZ-based approximate state preparation to achieve better fidelity. This work provides a path forward to benchmarking entanglement as NISQ devices improve in size and quality.

97 MATHEMATICS AND COMPUTING↗

A Quantum Annealing Computer Team Addresses Climate Change Predictability

The near confluence of the successful launch of the Orbiting Carbon Observatory2 on July 2, 2014 and the acceptance on August 20, 2015 by Google, NASA Ames Research Center and USRA of a 1152 qubit D-Wave 2X Quantum Annealing Computer (QAC), offered an exceptional opportunity to explore the potential of this technology to address the scientific prediction of global annual carbon uptake by land surface processes. At UMBC,we have collected and processed 20 months of global Level 2 light CO2 data as well as fluorescence data. In addition we have collected ARM data at 2sites in the US and Ameriflux data at more than 20 stations. J. Dorband has developed and implemented a multi-hidden layer Boltzmann Machine (BM) algorithm on the QAC. Employing the BM, we are calculating CO2 fluxes by training collocated OCO-2 level 2 CO2 data with ARM ground station tower data to infer to infer measured CO2 flux data. We generate CO2 fluxes with a regression analysis using these BM derived weights on the level 2 CO2 data for three Ameriflux sites distinct from the ARM stations. P. Gentine has negotiated for the access of K34 Ameriflux data in the Amazon and is applying a neural net to infer the CO2 fluxes. N. Talik validated the accuracy of the BM performance on the QAC against a restricted BM implementation on the IBM Softlayer Cloud with the Nvidia co-processors utilizing the same data sets. G. Nearing and K. Harrison have extended the GSFC LIS model with the NCAR Noah photosynthetic parameterization and have run a 10 year global prediction of the net ecosystem exchange. C. Pellisier is preparing a BM implementation of the Kalman filter data assimilation of CO2 fluxes. At UMBC, R. Prouty is conducting OSSE experiments with the LISNoah model on the IBM iDataPlex to simulate the impact of CO2 fluxes to improve the prediction of global annual carbon uptake. J. LeMoigne and D. Simpson have developed a neural net image registration system that will be used for MODIS ENVI and will be converted to a BM algorithm implementation on the QAC. The first integer adder has been implemented on the D-Wave 2X by A. Shehab that will perform HAAR wavelets for image compression of MODIS scenes. Finally, based on the next generations of QACs, we are preparing a 5-year performance road map on the scalability of the current QAC algorithms.

Science Data Processing↗

End-to-End Workflow for Machine-Learning-Based Qubit Readout With QICK and hls4ml

In this article, we present an end-to-end workflow for superconducting qubit readout that embeds codesigned neural networks into the quantum instrumentation control kit (QICK). Capitalizing on the custom firmware and software of the QICK platform, which is built on Xilinx radiofrequency system-on-chip field-programmable gate arrays (FPGAs), we aim to leverage machine learning (ML) to address critical challenges in qubit readout accuracy and scalability. The workflow utilizes the hls4ml package and employs quantization-aware training to translate ML models into hardware-efficient FPGA implementations via user-friendly Python application programming interfaces. We experimentally demonstrate the design, optimization, and integration of an ML algorithm for single transmon qubit readout, achieving 96% single-shot fidelity with a latency of 32.25 ns and less than 16% FPGA lookup table resource utilization. Our results offer the community an accessible workflow to advance ML-driven readout and adaptive control in quantum information processing applications.

42 ENGINEERING↗

Scalable molecular dynamics on CPU and GPU architectures with NAMD

NAMD is a molecular dynamics program designed for high-performance simulations of very large biological objects on CPU- and GPU-based architectures. NAMD offers scalable performance on petascale parallel supercomputers consisting of hundreds of thousands of cores, as well as on inexpensive commodity clusters commonly found in academic environments. It is written in C++ and leans on Charm++ parallel objects for optimal performance on low-latency architectures. NAMD is a versatile, multipurpose code that gathers state-of-the-art algorithms to carry out simulations in apt thermodynamic ensembles, using the widely popular CHARMM, AMBER, OPLS, and GROMOS biomolecular force fields. Here, we review the main features of NAMD that allow both equilibrium and enhanced-sampling molecular dynamics simulations with numerical efficiency. We describe the underlying concepts utilized by NAMD and their implementation, most notably for handling long-range electrostatics; controlling the temperature, pressure, and pH; applying external potentials on tailored grids; leveraging massively parallel resources in multiple-copy simulations; and hybrid quantum-mechanical/molecular-mechanical descriptions. We detail the variety of options offered by NAMD for enhanced-sampling simulations aimed at determining free-energy differences of either alchemical or geometrical transformations and outline their applicability to specific problems. Last, we discuss the roadmap for the development of NAMD and our current efforts toward achieving optimal performance on GPU-based architectures, for pushing back the limitations that have prevented biologically realistic billion-atom objects to be fruitfully simulated, and for making large-scale simulations less expensive and easier to set up, run, and analyze. NAMD is distributed free of charge with its source code at www.ks.uiuc.edu.

high-performance computing↗

A density-functional theory study of the Al/AlOx/Al tunnel junction

The aluminum oxide tunnel junction is a key component of the majority of superconducting quantum devices. For high-quality, reproducible, and scalably manufacturable qubits, the ability to fabricate Josephson junctions (JJs) with a targeted critical current and high uniformity is essential. In this study, we use first-principles modeling to assess fundamental aspects of the atomic structure of both amorphous and crystalline aluminum oxide tunnel junctions and relate the structure to predicted performance metrics. We use modified ab initio molecular dynamics to develop realistic models of the tunnel junction, from which interface roughness and local thickness fluctuations are analyzed in an unbiased manner by training a neural network to identify the boundary between metal and oxide. We show that the effective thickness of the insulating part of the junction can be different from the apparent physical thickness. We calculate the rate of Cooper pair tunneling for the atomically resolved electrostatic potential using direct numerical solution in 3D, which shows a channeling effect that impacts the junction critical current. The predicted critical current is a useful JJ design parameter that can be accessed from the ab initio calculations without fitting parameters. To assess the limits of uniformity and fabrication choices (e.g., oxidation vs epitaxy), we compare the amorphous junctions to crystalline models, which show order of magnitude more efficient tunneling compared to the amorphous case, underlining the connection between atomistic structure and Cooper pair tunneling efficiency. Further, this work provides a foundation for ab initio materials design and evaluation to help accelerate future development of improved tunnel junctions.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Evaluating the Limits of QAOA Parameter Transfer at High-Rounds on Sparse Ising Models With Geometrically Local Cubic Terms

The emergent practical applicability of the Quantum Approximate Optimization Algorithm (QAOA) for approximate combinatorial optimization is a subject of considerable interest. One of the primary limitations of QAOA is the task of finding a set of good parameters, which is usually done using a variational optimization loop. Parameter transfer, or parameter concentration, is a phenomenon where QAOA angles trained on problem instances that are self-similar tend to perform well for other problem instances from that similar class. This suggests a potentially highly efficient and scalable non-variational learning method for QAOA angle finding. In this work, we systematically study QAOA parameter transferability from small problem sizes (16 and 27 decision variables) onto large problem instances (up to 156 qubits) for heavy-hex graph Ising models with geometrically local higher order terms using the Julia based QAOA simulation tool \texttt{JuliQAOA} to perform classical angle finding for up to $49$ QAOA layers ($p$). Parameter transfer of the fixed angles is validated using a combination of full statevector, Projected Entangled Pair States (PEPS), Matrix Product State (MPS), and LOWESA numerical simulations. We find that the QAOA parameter transfer from single instances applied to other (unseen) problem instances does not in general provide monotonically improving performance as a function of $p$ - there are many cases where the performance temporarily decreases as a function of $p$ - but despite this the transferred angles have a general trend of improved expectation value as the QAOA depth increases, in many cases converging close to the true ground-state energy of the $100+$ qubit instances. We also sample the hardware-compatible Ising models using the ensemble of transfer-learned QAOA parameters on several superconducting qubit IBM Quantum processors with 127, 133, and 156 qubits. We find continuous solution quality improvement of the hardware-compatible QAOA circuits run on the IBM NISQ processors up to $p=5$ on \texttt{ibm\_fez}, up to $p=9$ on \texttt{ibm\_torino}, and up to $p=10$ on \texttt{ibm\_pittsburgh}.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

HydraGNN_Predictive_GFM_2024 - Ensemble of predictive graph foundation models for ground state atomistic materials modeling

We provide the ensemble of fifteen pre-trained graph foundation models (GFMs) for atomistic materials modeling applications. Each one of the fifteen GFMs has been trained on five open-source datasets that (once aggregated) amount to over 154 million atomistic structures, which cover over two-thirds of the natural elements of the periodic table and that comprises a broad set of organic and inorganic compounds. This vast set of atomistic structures comprises ground state configurations that are dynamically stable (i.e., equilibrated structures with atomic forces approximately close to zero values) as well as dynamically unstable structures (i.e., non-equilibrium structures with non-negligible non-zero values of atomic forces). The ensemble of datasets aggregated does NOT include excited states. The datasets have been curated to remove atomistic structures with spectral norm of the force tensor above 100 eV/angstrom. Moreover, a linear term of the energy was computed for each dataset using a linear regression model that uses the chemical concentration of each natural element as regressor. The linear term predicted by the linear regression model has been subtracted from each original energy value to perform a re-alignment of the energy values across different electronic structures approximation theories performed to generate the diverse multi-source, multi-fidelity datasets. The folder "ADIOS_files" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "ADIOS_files" directory contains 6 sub-directories named as follows: - ANI1x-v3.bp - MPTrj-v3.bp - OC2020-20M-v3.bp - OC2020-v3.bp - OC2022-v3.bp - qm7x-v3.bp Each sub-directory contains the pre-processed datasets converted in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used to the development, training, and performance testing of the ensemble go predictive graph foundation models. Each GFM was developed using HydraGNN (https://github.com/ORNL/HydraGNN) as underlying graph neural network (GNN) architecture. The multi-task learning (MTL) capability of HydraGNN was used to simultaneously train the GFMs on labeled values for direct predictions of energy (a total system property of an atomistic structure that measures the chemical stability) and atomic forces (an atomic level property of an atomistic structure that measures the dynamical stability). The hyper parameters of the GFM have been tuned using scalable hyperparameter optimization (HPO) algorithms implemented in the software DeepHyper (https://github.com/deephyper/deephyper). The pre-training of each HPO trial was performed using distributed data parallelism (DDP) to scale the training across 128 compute nodes of the exascale OLCF supercomputer Frontier. Each HPO trial was trained only for 10 epochs and an early stopping was performed to avoid wasting significant computational resources on GNN architectures that were clearly underperforming. For each HPO trial, the 'omnistat' tool developed by (AMD Research - Advanced Micro Device) was used to measure the total energy consumption in kWh. The ensemble of GFMs was obtained by selecting the fifteen best performing HPO trials. Four models have been selected for their clear advantage in accuracy, and these are the GFMs with IDs 229, 156, 147, 260. Additional eleven models have been selected based on judicious balance between accuracy and energy consumption needed for training, and these are the GFMs with IDs 165, 78, 137, 1, 175, 171, 181, 67, 179, 167, 351. Each selected GFM of the ensemble was continued to cumulate a total of at most 30 epochs. In some cases, the total number of epochs actually performed was les than 30 due to two combined factors: (1) the size of the GFM (i.e., the number of model parameters to train) and (2) the total wall-clock time for which the computational resources could be allocated on OLCF-Frontier. The "Ensemble_of_models" directory contains 15 sub-directories named as follows: - gfm_0.229 - gfm_0.156 - gfm_0.147 - gfm_0.260 - gfm_0.165 - gfm_0.78 - gfm_0.137 - gfm_0.1 - gfm_0.175 - gfm_0.171 - gfm_0.181 - gfm_0.67 - gfm_0.179 - gfm_0.167 - gfm_0.351 Each one of these sub-directories refers to one of the fifteen HPO trials that have been selected to continue the pre-training with at most 30 epochs. With each sub-directory associated with a specific HPO trial, the following files can be found: - config.json: file for argument parsing to develop and train an HydraGNN architecture - gfm_0.ID_epoch_N.pk: file with model parameters for HPO ID trial after N epochs of training The ensemble of fifteen GFM architectures was used for (1) ensemble averaging to stabilize the predictions of energy and atomic forces after pre-training for post-processing analysis and (2) ensemble uncertainty quantification (UQ). The code used to develop, pre-train, and load the pre-trained models for post-processing analysis is available on the ORNL-GitHub at the following link: https://github.com/ORNL/HydraGNN/tree/Predictive_GFM_2024

36 MATERIALS SCIENCE↗