Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “interactive HPC”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

88 records · Page 5

Photoelectron spectroscopy and computational investigations of the electronic structures and noncovalent interactions of cyclodextrin-closo-dodecaborate anion complexes x-CD·B12X122- (x = a, ß, y; X = H, F)

We report a joint negative ion photoelectron spectroscopy (NIPES) and computational study on the electronic structures and noncovalent interactions of a series of cyclodextrin-closo-dodecaborate dianion complexes, ?-CD·B12X122- (? = a, ß, ?; X = H, F). The measured vertical / adiabatic detachment energies (VDEs / ADEs) are 1.15/0.93, 3.55/3.20, 3.90/3.60, and 3.85/3.60 eV for B12H122- and its a-, ß-, ?-CD complexes, respectively; while the corresponding values are 1.90/1.70, 4.00/3.60, 4.33/3.95, and 4.30/3.85 eV for the X = F case. These results show that the inclusion of B12X122- into the CD cavities greatly increase the electronic stability of the dianions. The effect of electronic stabilization for ß-CD is roughly the same as for ?-CD, both being considerably stronger than that for a-CD. Density functional theory (DFT) based geometry optimization reveals that B12X122- are inserted into CDs increasingly deeper from a-CD to ?-CD. The calculated VDEs and ADEs agree with the experiments well, particularly, reproducing the electron binding energy (EBE) trends. The molecular orbital analyses indicate that the most loosely bound photodetached electrons origin from the guest B12X122- moieties. In addition to a shift of all signals to larger EBE, significant changes in the signal patterns are observed. At low EBE, this is due to the splitting of highly degenerate B12X122- orbitals, while at high EBE, photodetachment from CD oxygens contributes to the new bands. The guest B12X122- and host CD nocovalent, size-specific interaction based on the independent gradient model (IGM) and energy decomposition analysis (EDA), is dominated by electrostatic interactions. The analysis further unravels unambiguiously the existence of dihydrogen bonding and how it affects the total energy that stabilizes the host-guest complexes of CDs·B12H122- compared to the general hydrogen bonding interaction in CDs·B12F122-. This work clearly exhibits strong influences on the electronic structures of dodecaborates upon clustering with CDs, with both size (a-, ß-, ?-) and molecular (X = H or F) specificities, thus providing critical molecular-level information on the cyclodextrin-closo-dodecaborate interactions of interest to medical applications, e.g. Boron neutron capature therapy. The NIPES experiments done at PNNL were supported by the U.S. Department of Energy (DOE), Office of Science, Office of Basic Energy Sciences, Division of Chemical Sciences, Geosciences and Bioscience (X.-B.W.) and was performed at the EMSL, a national scientific user facility sponsored by DOE’s Office of Biological and Environmental Research and located at Pacific Northwest National Laboratory. H.S. and Z.S. acknowledge the funding support of National Natural Science Foundation of China (Nos.11727810, 61720106009 and 21603074), the Science and Technology Commission of Shanghai Municipality (Nos. 19JC1412200), and the Program of Introducing Talents of Discipline to Universities 111 project (B12024). Z. L. thanks the China Scholarship Council (CSC) for financial support. We acknowledge the ECNU Multifunctional Platform for Innovation (001) and HPC Research Computing Team for providing computational and storage resources. J.W. acknowledges support from the Alexander von Humboldt foundation (Feodor Lynen Fellowship and Rückkehrerstipendium), a Freigeist fellowship of the Volkswagenfoundation and Prof. Vladimir A. Azov for helpful discussions.

Li, Zhipeng↗

Massively Parallel Bayesian Model Calibration and Uncertainty Quantification with Applications to Nuclear Fuels and Materials

The U.S. Department of Energy (DOE)’s Nuclear Energy Advanced Modeling and Simulation (NEAMS) program aims to develop predictive capabilities by applying computational methods to the analysis and design of advanced reactor and fuel cycle systems. This program has been providing engineering-scale support for the development of BISON, a high-fidelity and high-resolution fuel performance tool. Fuel behavior in a nuclear reactor is governed by a complex network of mechanisms interacting with various other physics aspects in the reactor system. Any model developed to represent the fuel behavior will likely be idealized resulting in uncertainties in their predictions compared to the observed data. As such, this report was motivated by the need to identify the sources of uncertainties and quantify and propagate them through the fuel model outputs. Such quantification of uncertainties will establish a level of model trustworthiness, identify approaches to improve the model trustworthiness, and even guide optimal experiment design for maximal information gain. To accomplish the uncertainty quantification for computational models, this report has relied on the Bayesian framework which provides probabilistic treatment of models their inputs and outputs. The current state-of-the-art on performing Bayesian Uncertainty Quantification (UQ) for nuclear engineering models using High Performance Computing (HPC) resources have been reviewed. Implementation of capabilities for massively parallel Bayesian UQ in Multiphysics Object-Oriented Simulation Environment (MOOSE) is discussed. Several verification cases are discussed to verify the accuracy of the quantified uncertainties using the developed computational capabilities in MOOSE. Then, the problem of quantifying the uncertainties in TRI-Structural isOtropic (TRISO) fuel silver release is addressed. For the first time, the uncertainties arising from the TRISO Fission Gas Release (FGR) model due to model inadequacy and experimental noise are quantified. Also, the Bayesian capabilities are applied to the calibration of the MATPRO creep model, a widely used model in several fuel assessment cases. The impact of the prediction uncertainties in the MATPRO model on the fuel cladding behavior as part of the TRIBULATION assessment case (which is an integral effects case) is investigated. This report concludes with a discussion on the future work for the UQ for computational models.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Enabling Large-Scale Condensed-Phase Hybrid Density Functional Theory Based Ab Initio Molecular Dynamics. 1. Theory, Algorithm, and Performance

By including a fraction of exact exchange (EXX), hybrid functionals reduce the self-interaction error in semilocal density functional theory (DFT) and thereby furnish a more accurate and reliable description of the underlying electronic structure in systems throughout biology, chemistry, physics, and materials science. However, the high computational cost associated with the evaluation of all required EXX quantities has limited the applicability of hybrid DFT in the treatment of large molecules and complex condensed-phase materials. To overcome this limitation, we describe a linear-scaling approach that utilizes a local representation of the occupied orbitals (e.g., maximally localized Wannier functions (MLWFs)) to exploit the sparsity in the real-space evaluation of the quantum mechanical exchange interaction in finite-gap systems. In this work, we present a detailed description of the theoretical and algorithmic advances required to perform MLWF-based ab initio molecular dynamics (AIMD) simulations of large-scale condensed-phase systems of interest at the hybrid DFT level. We focus our theoretical discussion on the integration of this approach into the framework of Car–Parrinello AIMD, and highlight the central role played by the MLWF-product potential (i.e., the solution of Poisson’s equation for each corresponding MLWF-product density) in the evaluation of the EXX energy and wave function forces. We then provide a comprehensive description of the exx algorithm implemented in the open-source Quantum ESPRESSO program, which employs a hybrid MPI/OpenMP parallelization scheme to efficiently utilize the high-performance computing (HPC) resources available on current- and next-generation supercomputer architectures. Furthermore, this is followed by a critical assessment of the accuracy and parallel performance (e.g., strong and weak scaling) of this approach when AIMD simulations of liquid water are performed in the canonical (NVT) ensemble. With access to HPC resources, we demonstrate that exx enables hybrid DFT-based AIMD simulations of condensed-phase systems containing 500–1000 atoms (e.g., (H₂O)₂₅₆) with a wall time cost that is comparable to that of semilocal DFT. In doing so, exx takes us one step closer to routinely performing AIMD simulations of complex and large-scale condensed-phase systems for sufficiently long time scales at the hybrid DFT level of theory.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Onset of Fluidization in MP-PIC Simulations using MFIX-Exa

Fluidized bed reactors are used across a variety of industries, including for energy processes like pyrolysis that result in low-cost energy products. Design and scale-up of fluidized beds is de-risked by modeling and simulation, utilizing tools like NETL’s MFIX-Exa High-Performance Computing (HPC) code for reacting multiphase flow. This report summarizes an investigation into the breadth of problems to which MFIX-Exa may be applied, specifically with regard to low fluid velocities and the onset of fluidization. A simple fluidization study is conducted both experimentally and numerically for particles of interest, then reactor simulations are compared to cold flow experiments for uniform distributor plates. Approaches for modeling bubble caps are also presented.

discrete particle method↗

Educating HPC Users in the use of advanced computing technology

We examine a multi-modal approach to educating and training users of an advanced computing technology testbed at the Institute for Advanced Computational Science at Stony Brook University. Ookami provides researchers worldwide with access to 176 Fujitsu A64FX compute nodes, this being the same processor technology powering the Japanese Fugaku supercomputer, the fastest computer in the world since June 2020. However, achieving high-performance on this Arm-based, leadership computing technology requires that users be familiar with details of computer architecture, performance analysis and modeling, and high-performance programming models that are commonly omitted in introductory programming courses. Indeed, regardless of their seniority, many of the testbed users are surprisingly unfamiliar with basic concepts such as vectorization, pipelining, latency/bandwidth, roofline models, computing energy/power, threads, and non-uniform memory access. These same concepts also pervade mainstream x86 technologies, so this is of widespread concern. Due to the national/global nature of our user community that is also very diverse in both discipline and experience, the inability to offer formal classes, and our experience that most people do not tend to read online documentation or training materials in sufficient depth, we have consciously employed multiple approaches that heavily emphasize (online) personal interactions and transfer of skills. Online documentation has been organized around best-practices and FAQs; twice-weekly hackathons and office hours via Zoom enable deep dives by both the team and the user community with multiple broad benefits; a Slack channel provides both real time and archived answers and discussions; and workshops, training and webinars target community needs as they arise. Furthermore, the perspective that these tools are being used in an educational setting rather than just for project communication makes them more effective and contributes to community success.

A64FX↗

Scale-up Unlearnable Examples Learning with High-performance Computing

Recent advancements in AI models, like ChatGPT, are structured to retain user interactions, which could inadvertently include sensitive healthcare data. In the healthcare field, particularly when radiologists use AI-driven diagnostic tools hosted on online platforms, there is a risk that medical imaging data may be repurposed for future AI training without explicit consent, spotlighting critical privacy and intellectual property concerns around healthcare data usage. Addressing these privacy challenges, a novel approach known as Unlearnable Examples (UEs) has been introduced, aiming to make data unlearnable to deep learning models. A prominent method within this area, called Unlearnable Clustering (UC), has shown improved UE performance with larger batch sizes but was previously limited by computational resources (e.g., a single workstation). To push the boundaries of UE performance with theoretically unlimited resources, we scaled up UC learning across various datasets using Distributed Data Parallel (DDP) training on the Summit supercomputer. Our goal was to examine UE efficacy at high-performance computing (HPC) levels to prevent unauthorized learning and enhance data security, particularly exploring the impact of batch size on UE’s unlearnability. Utilizing the robust computational capabilities of the Summit, extensive experiments were conducted on diverse datasets such as Pets, MedMNist, Flowers, and Flowers102. Our findings reveal that both overly large and overly small batch sizes can lead to performance instability and affect accuracy. However, the relationship between batch size and unlearnability varied across datasets, highlighting the necessity for tailored batch size strategies to achieve optimal data protection. The use of Summit’s high-performance GPUs, along with the efficiency of the DDP framework, facilitated rapid updates of model parameters and consistent training across nodes. Our results underscore the critical role of selecting appropriate batch sizes based on the specific characteristics of each dataset to prevent learning and ensure data security in deep learning applications. The source code is publicly available at https: // github. com/ hrlblab/ UE_ HPC .

Zhu, Yanfan [Vanderbilt University, Nashville, TN,↗

Detector and Beamline Simulation for Next-Generation High Energy Physics Experiments

The success of high energy physics programs relies heavily on accurate detector simulations and beam interaction modeling. The increasingly complex detector geometries and beam dynamics require sophisticated techniques in order to meet the demands of current and future experiments. Common software tools used today are unable to fully utilize modern computational resources, while data-recording rates are often orders of magnitude larger than what can be produced via simulation. In this paper, we describe the state, current and future needs of high energy physics detector and beamline simulations and related challenges, and we propose a number of possible ways to address them.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Q-IRIS: The Evolution of the IRIS Task-Based Runtime to Enable Classical-Quantum Workflows

Extreme heterogeneity in emerging HPC systems are starting to include quantum accelerators, motivating runtimes that can coordinate between classical and quantum workloads. We present a proof-of-concept hybrid execution framework integrating the IRIS asynchronous task-based runtime with the XACC quantum programming framework via the Quantum Intermediate Representation Execution Engine (QIR-EE). IRIS orchestrates multiple programs written in the quantum intermediate representation (QIR) across heterogeneous backends (including multiple quantum simulators), enabling concurrent execution of classical and quantum tasks. Although not a performance study, we report measurable outcomes through the successful asynchronous scheduling and execution of multiple quantum workloads. To illustrate practical runtime implications, we decompose a four-qubit circuit into smaller subcircuits through a process known as quantum circuit cutting, reducing per-task quantum simulation load and demonstrating how task granularity can improve simulator throughput and reduce queueing behavior -- effects directly relevant to early quantum hardware environments. We conclude by outlining key challenges for scaling hybrid runtimes, including coordinated scheduling, classical-quantum interaction management, and support for diverse backend resources in heterogeneous systems.

Miniskar, Narasinga Rao [ORNL] (ORCID:000000018259↗

Understanding Lustre Internals. Second Edition

The Lustre file system has become a preferred storage resource for systems on the Top500 list, and it is often the file system of choice for small- to medium-sized HPC systems that require parallel shared access to data. Several resources exist to help users deploy and configure Lustre, but the same cannot be said for resources that explain the inner workings of the Lustre source code. A previous ORNL technical report entitled "Understanding Lustre Filesystem Internals" (ORNL/TM-2009/117) provided an excellent summary of Lustre subsystem operations. However, that report is over a decade old and is based on Lustre version 1.6. Since that report was published, Lustre has evolved significantly. Several subsystems underwent significant code changes and many new features have been added to the file system, bringing the current Lustre version up to 2.15.This report aims to document and explain the internal workings of the latest version of the Lustre file system. It will provide more complete and up-to-date information than the previous technical report and should serve as a foundational document for anyone interested in Lustre software development. Key data structures will be described along with the APIs used for interaction among the various Lustre subsystems. Although the Lustre software is constantly being developed, the details in this document should remain relevant for the forseeable future.

97 MATHEMATICS AND COMPUTING↗

SEAS Communication Engine: An Extensible, Flexible Wrapper for Co-Simulation Agents

When modeling and analyzing the power grid and other large scale systems, researchers often express scenarios as optimization problems and feed them into advanced software solvers. In order to allow multiple solvers to communicate with each other and share data from different domains, the National Renewable Energy Laboratory (NREL) and associated Department of Energy (DOE) labs have developed a software framework called the Hierarchical Engine for Large-scale Infrastructure Co-Simulation (HELICS). HELICS allows cosimulation via a collection of client libraries for different languages that can be called from the appropriate optimization software. However, these client libraries do not provide a higher level of abstraction beyond reading and writing data off of the shared HELICS bus. In this paper, we describe a new software library called the SEAS Communication Engine that exposes a higher-level API for running cosimulation problems. The SEAS Engine provides a class-based abstraction on top of the Python HELICS client, in order to allow users to implement their domain-specific cosimulations without needing to interact with core HELICS primitives. This will make adoption of HELICS and cosimulation in general easier, by exposing a simpler API. In the second part of the paper, we validate our library on a collection of different simulation examples, including the canonical IEEE 13 Bus Feeder. Lastly, we demonstrate using the SEAS Engine to directly call domain-specific code written in the Julia programming language. Our hope is that this will serve as a template for easily calling software in different programming languages via the SEAS Engine, thereby avoiding code duplication and complexity.

co-simulation↗

US Department of Energy, Office of Science, High-Performance Computing Facility: 2023 Operational Assessment Oak Ridge Leadership Computing Facility

The Oak Ridge Leadership Computing Facility (OLCF) was established to accelerate scientific discovery by providing world-leading computational performance and advanced data infrastructure. As a US Department of Energy (DOE) Office of Science user facility, the OLCF has managed the successful deployment and operation of a succession of leadership-class resources dedicated to open science. In addition to these resources, the OLCF staff continually strive to develop innovative processes and technologies, improve security, and empower users through effective allocation management and comprehensive user support and training. These efforts support the advancement of science by the OLCF users and benefit high-performance computing (HPC) facilities around the world. In calendar year (CY) 2023, the OLCF supported 1,676 users and 598 projects and exceeded all targets for user satisfaction. The facility received an average satisfaction score of 4.52 out of 5 on the annual user survey, and 94% of respondents reported a high satisfaction rate with the OLCF overall. Of the 3,619 user tickets submitted in CY 2023, OLCF staff resolved 97% within 3 business days. The facility opened Frontier to full scientific operations this year. Two projects conducted on Frontier received the Association for Computing Machinery (ACM) Gordon Bell Prize and the Gordon Bell Special Prize for Climate Modeling, and a third earned a nomination as a Gordon Bell Prize finalist. The facility’s previous flagship machine, Summit, gained new life and was extended through 2024 in part to help provide resources to the Integrated Research Infrastructure (IRI) projects and the National Artificial Intelligence Research Resource (NAIRR) pilot program. The facility instantiated an Advanced Computing Ecosystem testbed in part to support IRI workflows. OLCF made interactivity easier and more accessible to users than ever through tools like Jupyter notebooks and workflows.

97 MATHEMATICS AND COMPUTING↗

CACTI Radar b1 Processing: Corrections, Calibrations, and Processing Report

The U.S. Department of Energy’s (DOE) Atmospheric Radiation Measurement (ARM) user facility deployed a large number of instruments to a region nearby the Sierra de Córdobas mountains in Argentina as part of the Cloud, Aerosol, and Complex Terrain Interactions (CACTI) field campaign (1). During this campaign, four radars were installed at a site outside of Villa Yacanto as shown in Figure 1. As part of a post-campaign effort to improve the usability of these data, a significant activity was undertaken towards the calibration, correction, and improvement of the data quality of these radar datastreams. This process in ARM nomenclature is referred to as generating a “b1” datastream. While these “b1” standards may imply different corrections or standards for various ARM instruments, for radars it refers to a datastream that has been calibrated (and cross-calibrated), including a serious effort to deliver the highest-quality (well-characterized) data possible. This report details (i) the status/quality of the original “a1” (raw) data, (ii) the corrections and calibrations that are applied to generate these b1 datastreams, (iii) the details of the applied algorithms and how radar offset/calibration numbers were determined for the eventual corrections, and (iv) the new and flexible plug-in-based processing system designed during CACTI for radar b1 activities (current, future) that interfaces with ARM’s Data Integrator (ADI) and high-performance computing (HPC) system.

47 OTHER INSTRUMENTATION↗

Workflows Community Summit: Tightening the Integration between Computing Facilities and Scientific Workflows

Scientific workflows are used almost universally across science domains for solving complex and largescale computing and data analysis problems. The importance of workflows is highlighted by the fact that they have underpinned some of the most significant discoveries of the past decades. Many of these workflows have significant computational, storage, and communication demands, and thus must execute on a range of large-scale computer systems, from local clusters to public clouds and upcoming exascale HPC platforms. Managing these executions is often a significant undertaking, requiring a sophisticated and versatile software infrastructure. Historically, infrastructures for workflow execution consisted of complex, integrated systems, developed in-house by workflow practitioners with strong dependencies on a range of legacy technologies—even including sets of ad hoc scripts. Due to the increasing need to support workflows, dedicated workflow systems were developed to provide abstractions for creating, executing, and adapting workflows conveniently and efficiently while ensuring portability. While these efforts are all worthwhile individually, there are now hundreds of independent workflow systems. These workflow systems are created and used by thousands of researchers and developers, leading to a rapidly growing corpus of workflows research publications. The resulting workflow system technology landscape is fragmented, which may present significant barriers for future workflow users due to many seemingly comparable, yet usually mutually incompatible, systems that exist. In order to tackle some of the challenges described above, the DOE-funded ExaWorks and NSF-funded WorkflowsRI projects have organized in 2021 a series of events entitled the “Workflows Community Summit”. The third edition of the “Workflows Community Summit” explored workflows challenges and opportunities from the perspective of computing centers and facilities. This third summit builds on two prior summits (https://workflowsri.org/summits) that (i) established a high level vision for workflows research; and (ii) explored technical approaches for realizing that vision. The third summit brought together a small group of facilities representatives with the aim to understand how workflows are currently being used at each facility, how facilities would like to interact with workflow developers and users, how workflows fit with facility roadmaps, and what opportunities there are for tighter integration between facilities and workflows. This report documents and organizes the wealth of information provided by the participants before, during, and after the summit.

97 MATHEMATICS AND COMPUTING↗

Improved Distributed-memory Triangle Counting by Exploiting the Graph Structure

Graphs are ubiquitous in modeling complex systems and representing interactions between entities to uncover structural information of the domain. Traditionally, graph analytics workloads are challenging to efficiently scale (both strong and weak cases) on distributed memory due to the irregular memory-access driven nature (with little or no computations) of the methods. The structure of graphs and their relative distribution over the processing elements poses another level of complexity, making it difficult to attain sustainable scalability across platforms. In this paper, we discuss enhancements to TriC, a distributed-memory implementation of graph triangle counting using Message Passing Interface (MPI), which was featured in the 2020 Graph Challenge competition. We have made some incremental enhancements to TriC, primarily adopting a user-defined buffering strategy to overcome the startup problem for large graphs (by fixing the memory for intermediate data), and experimenting with probabilistic data structures such as bloom filter to improve the query response time for assessing edge existence, at the expense of increasing the overall false positive rate. These adjustments have led to a modest improvements in most cases, as compared to the previous version.

Graph Analytics, HPC↗

Modeling the Interaction of Laser-Produced Proton Beams with Matter

A major goal of this project is to significantly increase our understanding of isochoric heating of matter using laser produced proton beams, and the associated high energy density (HED) and warm dense matter (WDM) regimes generated. This will benefit research fields such as planetary science, fusion energy, plasma physics, and material science. For example, it will enhance our understanding of WDM properties of iron and silica under conditions encountered in planetary interiors and diagnostic components in fusion devices exposed to high fluxes of energetic plasma ions. The project is motivated by recent experiments that irradiated Si targets with proton beams generated by the 20 TW-laser at the SLAC MEC end-station. The HED/WDM states are probed using the 50 fs hard X-rays available in the 3rd harmonic of the LCLS. As part of this project, results from the phase contrast X-ray imaging, which shows the generation of compression waves that produces rear surface spallation, are compared with results from the 3D multi-physics multi- material code, PISALE, that combines Arbitrary Lagrangian-Eulerian (ALE) hydrodynamics with Adaptive Mesh Refinement (AMR). This comparison required modifications to several physics models in the PISALE (Pacific Island Structured-AMR with ALE) code. An important aspect of this project is the continued training of graduate students in HED physics and in conducting complex multiphysics simulations.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

matsim-agents v1.0

matsim-agents is a multi-agent AI framework for atomistic materials simulation and discovery. It orchestrates large language models (LLMs), machine-learned interatomic potentials (MLIPs), and DFT codes into a single agentic loop running on laptops and DOE leadership-class supercomputers. MULTI-AGENT ORCHESTRATION A LangGraph state machine with three nodes: a Planner that converts a natural-language research objective into structured tasks; an Executor that dispatches atomistic tools and loops until the queue is empty; and an Analyst that summarizes results into a human-readable report. State is checkpointed after every step and human-in-the-loop gates can be inserted at any edge. HYPOTHESIS-DRIVEN DISCOVERY CHAT An interactive REPL (matsim-agents chat) that couples LLM dialogue with atomistic simulation. Chemical formulas are automatically detected in conversation turns and trigger a full crystal-phase exploration: structure generation → relaxation → stability scoring → result injection back into the conversation, creating a closed hypothesis-refinement loop. CRYSTAL PHASE ENUMERATION Given a composition, the phase explorer enumerates prototypes by stoichiometry: elemental (fcc/bcc/hcp/sc/diamond), binary 1:1 (rocksalt/CsCl/zincblende/ wurtzite/fluorite/rutile), ternary 1:1:3 (cubic perovskite), ternary 1:2:4 (perovskite + spinel), quaternary 1:1:2:6 (Fm-3m double perovskite). 2-D prototypes (graphene, h-BN, MoS2 2H/1T) and multilayer stacking are also supported via --include-2d and --num-layers. SUPERCELL GENERATION AND SITE DECORATION Auto-tiling to a minimum atom count (--min-atoms), explicit NxNxN tiling (--supercell), symmetry-distinct site decorations (--n-orderings), and isotropic lattice-scale sweeps (--lattice-scales) for volume bracketing. MLFF RELAXATION AND STABILITY SCORING HydraGNN (multi-headed GNN) drives structure relaxation via ASE with FIRE, BFGS, or BFGSLineSearch. Stability output: delta-E/atom ranking across phases and a max-residual-force dynamical-stability proxy. Other MLIPs (MACE, NequIP, Orb) can be plugged in through the same interface. DFT BACKENDS Quantum ESPRESSO pw.x and VASP 6.6 are first-class labellers. Both have validated GPU builds and SLURM/PBS launchers for three DOE platforms: Frontier (AMD MI250X, ROCm), Aurora (Intel PVC, oneAPI), Perlmutter (NVIDIA A100, CUDA). QE produces ~100 binaries (pw.x, ph.x, epw.x, ...). VASP supports scf, relax, vc-relax, and vc-relax-shape run types. ACTIVE-LEARNING LOOP matsim-agents al run CONFIG.yaml drives an iterative HydraGNN-DFT loop: MD generates candidates → ensemble/MC-dropout uncertainty selects the most informative → DFT labels them in parallel inside one allocation → dataset grows → HydraGNN retrains → repeat. DFT backend is a single YAML toggle (dft.backend: vasp | qe). LLM-generated seed structures are supported (no curated POSCAR library needed). Config uses ${VAR}, ${VAR:-default}, ${VAR:?msg} shell-style substitution for cross-user/cross-site portability. LLM BACKENDS Ollama (local, default), vLLM (HPC multi-GPU serving), OpenAI, Anthropic, HuggingFace Transformers+Accelerate. Selected at runtime via flag or env var with no code changes. HPC PORTABILITY Same Python entry points run on Frontier (ROCm 7.2), Aurora (oneAPI), and Perlmutter (CUDA 12). DFT and ML stacks are never co-loaded in the same shell; they couple through the scheduler and filesystem. Advanced multi-node launchers (serve, discovery-chat, single-relaxation, active-learning, QE warm-start) are provided for all three platforms. CODABENCH COMPETITION BUNDLE A self-contained benchmark: 159 atomistic test structures across 11 material classes, 5 tasks (formation energy, forces, ML relaxation, AI-DFT relaxation, phase stability ranking), public/private leaderboard split (30/70), and four ready-to-run baselines: MACE-MP-0, HydraGNN, UMA, AllScAIP.

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗