Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Architecture patterns”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Adapting In Situ Accelerators for Sparsity With Granular Matrix Reordering

Neural network (NN) inference is an essential part of modern systems and is found at the heart of numerous applications ranging from image recognition to natural language processing. In situ NN accelerators can efficiently perform NN inference using resistive crossbars, which makes them a promising solution to the data movement challenges faced by conventional architectures. Although such accelerators demonstrate significant potential for dense NNs, they often do not benefit from sparse NNs, which contain relatively few non-zero weights. Processing sparse NNs on in situ accelerators results in wasted energy to charge the entire crossbar where most elements are zeros. To address this limitation, this paper proposes Granular Matrix Reordering (GMR): a preprocessing technique that enables an energy-efficient computation of sparse NNs on in situ accelerators. GMR reorders the rows and columns of sparse weight matrices to maximize the crossbars' utilization and minimize the total number of crossbars needed to be charged. The reordering process does not rely on sparsity patterns and incurs no accuracy loss. Finally, GMR achieves an average of 28% and up to 34% reduction in energy consumption over seven pruned NNs across four different pruning methods and network architectures.

97 MATHEMATICS AND COMPUTING↗

Genetic architecture of leaf morphological and physiological traits in a Populus deltoides ‘Danhong’ × P. simonii ‘Tongliao1’ pedigree revealed by quantitative trait locus analysis

Understanding the genetic architecture of leaf morphological and physiological traits will help plant breeders develop high biomass poplar genotypes. Quantitative trait locus (QTL) studies combining next-generation sequencing techniques can advance our understanding of the genetic basis of complex traits. In this study, we measured 13 leaf morphological and physiological traits and identified quantitative trait loci (QTLs) in a Populus deltoides ‘Danhong’ × P. simonii ‘Tongliao1’ F1 population (500 progenies) using a high-density genetic map constructed by whole genome re-sequencing. This linkage map consisted of 5796 single nucleotide polymorphism (SNP) markers assigned to 19 linkage groups (LGs), spanning 2683.80 centimorgans (cM) of genetic length, with an average marker density of 0.46 cM. We identified 109 QTLs on 18 LGs for leaf morphological traits and 55 QTLs on 14 LGs for leaf physiological traits. One-hundred eight putative candidate genes were identified within the candidate genomic region. Co-expression network and gene ontology enrichment analyses suggested that these candidate genes were involved in the photosynthetic process. The differential expression patterns of the CYCLIN (Potri.015G112200) and RED CHLOROPHYLL REDUCTASE (Potri.007G043600) genes between two parents indicated their potential roles in leaf morphological and physiological traits. These findings decipher the genetic architecture of leaf morphological and physiological traits in the P. deltoides ‘Danhong’ × P. simonii ‘Tongliao1’ pedigree and provide candidate genes for future poplar genetic improvement.

59 BASIC BIOLOGICAL SCIENCES↗

BULKI-Store v0.3.2

BULKI-Store is a distributed object storage system optimized for high-performance computing environments. Built with a Rust core and Python bindings, it efficiently manages scientific and machine learning datasets across HPC clusters. The system employs a client-server architecture with MPI integration, enabling seamless scaling on supercomputers like Perlmutter. BULKI-Store's object-oriented approach provides intuitive data organization with rich metadata support, contrasting with traditional file-based solutions. Key optimizations include selective checkpoint loading, unified checkpoint files, and object chunking for large data transfers. For machine learning workloads, BULKI-Store offers advantages through fine-grained access patterns, dynamic data sharing between training instances, and reduced memory pressure. Memory management features include strategic Python GC calls, minimized data copies, and batch processing capabilities. The system leverages Rayon's thread pool for asynchronous data prefetching and supports multiple CPU architectures (ARM64, x86, AMD, RISC-V). By combining performance optimizations with developer-friendly APIs, BULKI-Store addresses the complex data management challenges of modern HPC applications while maintaining compatibility across heterogeneous computing environments.

Zhang, Wei [Lawrence Berkeley National Laboratory ↗

RingX: Scalable Parallel Attention for Long-Context Learning on HPC

The attention mechanism has become foundational for remarkable AI breakthroughs since the introduction of the Transformer, driving the demand for increasingly longer context to power frontier models such as large-scale reasoning language models and high-resolution image/video generators. However, its quadratic computational and memory complexities present substantial challenges. Current state-of-the-art parallel attention methods, such as ring attention, are widely adopted for long-context training but utilize a point-to-point communication strategy that fails to fully exploit the capabilities of modern HPC network architectures. In this work, we propose ringX, a scalable family of parallel attention methods optimized explicitly for HPC systems. By enhancing workload partitioning, refining communication patterns, and improving load balancing, ringX achieves up to 3.4 × speedup compared to conventional ring attention on the Frontier supercomputer. Optimized for both bi-directional and causal attention mechanisms, ringX demonstrates its effectiveness through training benchmarks of a Vision Transformer (ViT) on a climate dataset and a Generative Pre-Trained Transformer (GPT) model, Llama3 8B. Our method attains an end-to-end training speedup of approximately 1.5 × in both scenarios. To our knowledge, the achieved 38% model FLOPs utilization (MFU) for training Llama3 8B with a 1M-token sequence length on 4,096 GPUs represents one of the highest training efficiencies reported for long-context learning on HPC systems. Our code implementation is available at https://github.com/jqyin/ringX-attention.

Yin, Junqi [ORNL] (ORCID:0000000338435520)↗

Sensing Electrical Networks Securely & Economically (SENSE)

The growing adoption of distributed energy resources (DERs) like battery energy storage systems and roof top solar/PV and the rapid penetration of electric vehicles (EVs), the electric grid is undergoing a major transformation with elevated stress on legacy grid assets. Despite a lot of expenditure to address these challenges, both in dollars and manpower, utilities have not been able to receive the value that was promised. The gains have been most visible at the transmission and substation level, especially where the main objective was improving operational and economic efficiency for the utility. Improving visibility and control at a few select points enhances the existing and established paradigm of centralized command and control. With changing load patterns, load types and the overall transition to an “active grid”, the centralized control and coordination paradigm gets challenged. To address the challenges, a new architecture and mechanism is needed, one that supports decentralized control and decision making, extracting value streams at the grid edge, particularly as the changes are fueled by transitions occurring in the distribution system. To address this, a communications and data processing platform, “GAMMA” was developed and demonstrated through the project. At the heart of the platform, are distributed, intelligent edge nodes with sensing and compute capabilities, that can record and analyze information locally. They are embedded in sensors and actuators specific to different distribution system applications. Phase 1 of the project focused on developing novel sensor technology that can be used for monitoring utility pole top distribution transformers. The sensors were designed with the objective of being low-cost, communicating with the GAMMA cloud using novel “delay-tolerant” networking using Bluetooth and a secure mobile application. They were non-intrusive in nature so that they can be installed quickly in the field, resulting in overall low cost of deployment and operations. Following the successful completion of Phase 1, the team manufactured 100 units for a field demonstration in Phase 2. The field demonstration was carried out on two real feeder systems with the local utility partner. In total, 100 sensors were installed and operated over a period of 6 months in the state of Georgia. The platform is operational end to end, with the cloud infrastructure deployed on a distributed, serverless environment that can serve multiple data streams, an analytics engine and a portal to securely view the data from multiple assets. The data collected through the GAMMA Mobile Phone app showcased the viability of the novel delay tolerant networking architecture, and the data processing algorithms developed through the course of the project, were successful in extracting important information about the overall network, improving the utility’s visibility and situational awareness in the distribution feeder.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Kinetic Inductance Phonon-Mediated Detectors for Dark Matter: R&D at NEXUS

Superconducting thin films have long been used as phonon sensors, particularly in the field of dark matter (DM) direct detection, due to the meV-scale Cooper-pair binding energy. A novel class of these detectors based on microwave kinetic inductance detectors, dubbed Kinetic Inductance Phonon-Mediated Detectors (KIPMDs), offers an attractive architecture for microcalorimeters to probe DM down to the fermionic thermal relic mass limit of a few keV. Such a device featuring an aluminum resonator patterned onto a silicon substrate was operated at the NEXUS low-background facility at Fermilab for characterization and evaluation of its efficacy for a dark matter search. With this device we have demonstrated a resolution on the energy absorbed by the superconductor of 2.5 eV, a factor of two better than current state-of-the-art. In this talk, I will present our measurement of the energy resolution and phonon collection efficiency performed by exposing the bare substrate to a pulsed source of 470 nm photons. I will also discuss the path forward to obtaining in these devices the sub-eV resolution required to test the “Freeze-in” class of DM models. Finally, I will review other complementary efforts in our group to develop a superconducting qubit-based low-threshold detector.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Designing Energy-Efficient Quantum Computers Through Prediction and Reduction of Cooling Requirements for Cryogenic Electronics

Quantum computing has been identified as a “wild card” by the International Energy Agency in predicting future global data center energy usage. This is primarily because both uncertainty in the extent to which quantum computing will be adopted, and uncertainty in the power consumption of individual quantum data centers. Unlike the classical counterparts, quantum computers need to be maintained at near absolute zero, requiring energy-intensive cryogenic cooling systems. Therefore, as quantum computers scale up from existing 50 qubit technology demonstrations to the 10,000 to 100,000 qubit systems that will be able to solve complex problems, the energy consumption of both the electronics and the required cooling systems will also increase. To predict this scaling, this work analyzes the energy requirements for both computation and cooling of quantum hardware. We show that the energy requirements for cooling of quantum computers is determined by several computing system parameters, including the number and type of physical qubits, the operating temperature, the packaging efficiency of the system, and the split between circuits operating at cryogenic temperatures and those operating at room temperature. The energy requirements can then be found based on thermal system parameters such as cooling efficiency and cryostat heat transfer. Analysis of these parameters shows that the energy required for cooling is significantly larger than that required for computation, a reversal from energy usage patterns seen in conventional computing. The results and discussions provide a road-map for creating energy efficient quantum computers through the selection of computer architectures and cryogenic system configurations that minimize cooling requirements.

energy efficiency↗

Scalable Pattern Matching in Metadata Graphs via Constraint Checking

Pattern matching is a fundamental tool for answering complex graph queries. Unfortunately, existing solutions have limited capabilities: They do not scale to process large graphs and/or support only a restricted set of search templates or usage scenarios. Moreover, the algorithms at the core of the existing techniques are not suitable for today’s graph processing infrastructures relying on horizontal scalability and shared-nothing clusters, as most of these algorithms are inherently sequential and difficult to parallelize. In this article we present an algorithmic pipeline that bases pattern matching on constraint checking. The key intuition is that each vertex and edge participating in a match has to meet a set of constraints implicitly specified by the search template. These constraints can be verified independently and typically are less expensive to compute than searching the full template. The pipeline we propose generates these constraints and iterates over them to eliminate all the vertices and edges that do not participate in any match, thus reducing the background graph to a subgraph that is the union of all template matches—the complete set of all vertices and edges that participate in at least one match. Additional analysis can be performed on this annotated, reduced graph, such as full match enumeration, match counting, or computing vertex/edge centrality. Furthermore, a vertex-centric formulation for constraint checking algorithms exists, and this makes it possible to harness existing high-performance, vertex-centric graph processing frameworks. This technique (i) enables highly scalable pattern matching in metadata (labeled) graphs; (ii) supports arbitrary patterns with 100% precision; (iii) enables tradeoffs between precision and time-to-solution, while always selects all vertices and edges that participate in matches, thus offering 100% recall; and (iv) supports a set of popular data analytics scenarios. We implement our approach on top of HavoqGT, an open-source asynchronous graph processing framework, and demonstrate its advantages through strong and weak scaling experiments on massive scale real-world (up to 257 billion edges) and synthetic (up to 4.4 trillion edges) labeled graphs, respectively, and at scales (1,024 nodes / 36,864 cores), orders of magnitude larger than used in the past for similar problems. This article serves two purposes: First, it synthesises the knowledge accumulated during a long-term project. Second, it presents new system features, usage scenarios, optimizations, and comparisons with related work that strengthen the confidence that pattern matching based on iterative pruning via constraint checking is an effective and scalable approach in practice. The new contributions include the following: (i) We demonstrate the ability of the constraint checking approach to efficiently support two additional search scenarios that often emerge in practice, interactive incremental search and exploratory search. (ii) We empirically compare our solution with two additional state-of-the-art systems, Arabsque and TriAD. (iii) We show the ability of our solution to accommodate a more diverse range of datasets with varying properties, e.g., scale, skewness, label distribution, and match frequency. (iv) We introduce or extend a number of system features (e.g., work aggregation, load balancing, and the ability to cap the generated traffic) and design optimizations and demonstrate their advantages with respect to improving performance and scalability. (v) We present bottleneck analysis and insights into artifacts that influence performance. (vi) We present a theoretical complexity argument that motivates the performance gains we observe.

97 MATHEMATICS AND COMPUTING↗

Ensuring reliable connectivity to cellular-connected UAVs with up-tilted antennas and interference coordination

To integrate unmanned aerial vehicles (UAVs) in future large-scale deployments, a new wireless communication paradigm, namely, the cellular-connected UAV has recently attracted interest. However, the line-of-sight dominant air-to-ground channels along with the antenna pattern of the cellular ground base stations (GBSs) introduce critical interference issues in cellular-connected UAV communications. In particular, the complex antenna pattern and the ground reflection (GR) from the down-tilted antennas create both coverage holes and patchy coverage for the UAVs in the sky, which leads to unreliable connectivity from the underlying cellular network. To overcome these challenges, in this paper, we propose a new cellular architecture that employs an extra set of co-channel antennas oriented towards the sky to support UAVs on top of the existing down-tilted antennas for ground user equipment (GUE). To model the GR stemming from the down-tilted antennas, we propose a path-loss model, which takes both antenna radiation pattern and configuration into account. Next, we formulate an optimization problem to maximize the minimum signal-to-interference ratio (SIR) of the UAVs by tuning the up-tilt (UT) angles of the up-tilted antennas. Since this is an NP-hard problem, we propose a genetic algorithm (GA) based heuristic method to optimize the UT angles of these antennas. After obtaining the optimal UT angles, we integrate the 3GPP Release-10 specified enhanced inter-cell interference coordination (eICIC) to reduce the interference stemming from the down-tilted antennas. Our simulation results based on the hexagonal cell layout show that the proposed interference mitigation method can ensure higher minimum SIRs for the UAVs over baseline methods while creating minimal impact on the SIR of GUEs.

3GPP↗

I/O Access Patterns in HPC Applications: A 360-Degree Survey

The high-performance computing I/O stack has been complex due to multiple software layers, the inter-dependencies among these layers, and the different performance tuning options for each layer. In this complex stack, the definition of an “I/O access pattern” has been reappropriated to describe what an application is doing to write or read data from the perspective of different layers of the stack, often comprising a different set of features. It has become common to have to redefine what is meant when discussing a pattern in every new study, as no assumption can be made. This survey aims to propose a baseline taxonomy, harnessing the I/O community’s knowledge over the past 20 years. This definition can serve as a common ground for high-performance computing I/O researchers and developers to apply known I/O tuning strategies and design new strategies for improving I/O performance. We seek to summarize and bring a consensus to the multiple ways to describe a pattern based on common features already used by the community over the years.

97 MATHEMATICS AND COMPUTING↗

Demystifying asynchronous I/O Interference in HPC applications

With increasing complexity of HPC workflows, data management services need to perform expensive I/O operations asynchronously in the background, aiming to overlap the I/O with the application runtime. However, this may cause interference due to competition for resources: CPU, memory/network bandwidth. The advent of multi-core architectures has exacerbated this problem, as many I/O operations are issued concurrently, thereby competing not only with the application but also among themselves. Furthermore, the interference patterns can dynamically change as a response to variations in application behavior and I/O subsystems (e.g. multiple users sharing a parallel file system). Without a thorough understanding, I/O operations may perform suboptimally, potentially even worse than in the blocking case. To fill this gap, here we investigate the causes and consequences of interference due to asynchronous I/O on HPC systems. Specifically, we focus on multi-core CPUs and memory bandwidth, isolating the interference due to each resource. Then, we perform an in-depth study to explain the interplay and contention in a variety of resource sharing scenarios such as varying priority and number of background I/O threads and different I/O strategies: sendfile, read/write, mmap/write underlining trade-offs. The insights from this study are important both to enable guided optimizations of existing background I/O, as well as to open new opportunities to design advanced asynchronous I/O strategies.

97 MATHEMATICS AND COMPUTING↗

Engineering catalytical electrodes for applications in energy areas

An ink formulation and electrode that enhances hydrogen production, oxygen production, carbon dioxide reduction and other electrocatalytic reactions. Embodiments include an ink formulation with polymer binders having different catalytical precursors and a 3D electrode produced by additive manufacturing from the inventor's ink formulation. Various embodiments of the inventor's apparatus, systems, and methods provide inks that that are 3D-printed into patterns that optimize surface area and flow. The catalytic materials are imbedded into the ink matrix which is then printed into a 3D structure that has architecture that optimizes surface area and flow properties.

Liang, Siwei↗

Resilient Information Architecture Platform for Smart Grid (RIAPS)

A number of emerging trends will substantially alter the operation and control of the electric grid over the next several decades. These trends include ensuring resiliency under severe weather events, increasing integration of renewable electricity generation, supporting changing electricity demand patterns, and the improving cost effectiveness of distributed energy resources. To address these challenges, the future “Smart Grid” management will need to transition from centralized to coordinated distributed control paradigm. Reliable operation of the Smart Grid depends on distributed intelligence realized through software applications that run on distributed computing devices attached to the power system to collect data and collaboratively manage resources. However, much of the existing software for Smart Grid-enabled devices is either proprietary or developed with custom solutions, which limits interoperability among the heterogeneous devices and hinders the ability to manage system-level reliability, security, and resiliency requirements. Additionally, this approach makes Smart Grid applications hard to maintain, evolve, verify, and replace; resulting in high development and deployment costs. Further development of the Smart Grid requires a reusable software base-layer to move from hard-coded functionality to a plug-and-play architecture capable of managing system-level objectives and constraints in addition to providing consistent common services across heterogeneous devices and applications. Vanderbilt University, in collaboration with North Carolina State University and Washington State University has developed a foundation ‘software platform’ for developing and deploying robust, reliable, effective and secure software applications for the Smart Grid. The Resilient Information Architecture Platform for the Smart Grid (RIAPS) provides core services for building effective and powerful smart grid applications. It offers unique services for real-time data dissemination, fault tolerance, and coordination across apps distributed over the network.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Explore Spatio‐Temporal Learning of Large Sample Hydrology Using Graph Neural Networks

Abstract Streamflow forecasting over gauged and ungauged basins play a vital role in water resources planning, especially under the changing climate. Increased availability of large sample hydrology data sets, together with recent advances in deep learning techniques, has presented new opportunities to explore temporal and spatial patterns in hydrological signatures for improving streamflow forecasting. The purpose of this study is to adapt and benchmark several state‐of‐the‐art graph neural network (GNN) architectures, including ChebNet, Graph Convolutional Network (GCN), and GraphWaveNet, for end‐to‐end graph learning. We explicitly represent river basins as nodes in a graph, learn the spatiotemporal nodal dependencies, and then use the learned relations to predict streamflow simultaneously across all nodes in the graph. The efficacy of the developed GNN models is investigated using the Catchment Attributes and MEteorology for Large‐sample Studies (CAMELS) data set under two settings, fixed graph topology (transductive learning), and variable graph topology (inductive learning), with the latter applicable to prediction in ungauged basins (PUB). Results indicate that GNNs are generally robust and computationally efficient, achieving similar or better performance than a baseline model trained using the long short‐term memory (LSTM) network. Further analyses are conducted to interpret the graph learning process at the edge and node levels and to investigate the effect of different model configurations. We conclude that graph learning constitutes a viable machine learning‐based method for aggregating spatiotemporal information from a multitude of sources for streamflow forecasting

Sun, Alexander Y.↗

DOE Repository Metadata Profile (DRMP): A Metadata Framework for Advancing Interoperability and AI Readiness Across Scientific Repositories

The Department of Energy (DOE) funds a diverse and distributed ecosystem of repositories that steward scientific data, publications, and software across its research programs, user facilities, and national laboratories. While significant progress has been made in standardizing dataset-level metadata, the metadata describing repositories themselves (their identity, governance, access interfaces, policies, and technical capabilities) remains inconsistent and fragmented across DOE-funded systems. This variability limits discoverability, interoperability, automated validation, and AI-driven analysis, all of which are increasingly essential for modern scientific workflows. To address this gap, the DOE Data Curation Working Group (DCWG) developed the DOE Repository Metadata Profile (DRMP). The DRMP is a practical, community-driven framework that defines how repositories can describe themselves in a consistent, machine-actionable, and scalable manner. The DRMP is not a new metadata schema. Instead, it is a mapping profile and structured element set capturing the essential characteristics of DOE repositories. It harmonizes repository-level metadata across six widely adopted community schemas: RE3Data; DCAT-US v3; Schema.org; Dublin Core; DataCite 4.6; and PREMIS 3.0. This harmonization eliminates reinvention and enables interoperability within DOE and across the broader scientific ecosystem. A core objective of the DRMP is to reduce burden on repositories by allowing them to reuse their existing metadata through a Rosetta-style crosswalk rather than redesigning local implementations. The profile introduces a three-level conformance model that supports incremental adoption: • Level 1 – Minimum Viable Record (MVR): foundational identification elements required for workflows, project registration, and basic repository presence. • Level 2 – Interoperable: structured metadata enabling alignment with national and international discovery systems. • Level 3 – AI-Ready: enhanced provenance, policy transparency, fixity, semantic context, and capabilities that support automated reasoning, model training governance, and machine-assisted curation. To support implementation, the DRMP includes JSON Schema definitions, OpenAPI patterns, and MCP templates that allow repositories to publish machine-readable metadata directly within existing platforms. These resources are modular and lightweight, enabling adoption without major architectural change. Adopting the DRMP enables repositories to: • Enhance discoverability and interoperability by aligning identifiers, classifications, and descriptive elements across widely used schema standards. • Support federated discovery and cross-registration across DOE systems, Data.gov, and international catalogs. • Enable AI agents and workflow orchestration systems to interpret repository-level metadata within the American Science Cloud (AmSC) through Model Context Protocol (MCP)-based context publication. • Demonstrate alignment with DOE’s open science, stewardship, and FAIR data priorities. This guidance represents a community-driven step forward. Through voluntary adoption and continued feedback, the DRMP advances a cohesive, machine-actionable description of DOE repositories that supports FAIR data practices, preparing the infrastructure for AI-enabled research, and strengthening the discoverability and reuse of DOE’s scientific outputs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Evaluation of Graph Analytics Frameworks Using the GAP Benchmark Suite

The analysis of connected data is an increasingly important application in high-performance computing. Such analyses can reveal fraudulent patterns in financial transactions, optimize telecommunications networks, predict information flow in social networks, etc. However, the landscape of graph analytics is highly diverse. Graph algorithms stress processor architectures differently, and no one graph can represent all topologies. Consequently, no single approach or framework is expected to be optimal for all graph analytics problems. To help make sense of this diverse landscape, we evaluated four approaches to graph analytics: GraphBLAS, Galois, BGL17, GraphIt; and compare them against hand-tuned implementations that take advantage of hardware features on our test platform. Graph- BLAS formulates graph analytics as sparse linear algebra. Galois provides syntactic constructs for data parallelism over irregular data structures. BGL17 is a generic C++ template library for implementing graph algorithms. GraphIt provides a domain- specific language to describe and optimize graph algorithms. We use the GAP Benchmark Suite to establish baseline performance and guide the side-by-side evaluation of each framework. GAP consists of 30 tests: six graph analytics algorithms (breadth- first search, single-source shortest path, PageRank, betweenness centrality, connected components, and triangle counting) run on five graphs, each with different topological characteristics (e.g., high diameter, skewed degree distribution, high average degree). High-performance reference implementations are included for each benchmark algorithm. Because a graph can be loaded into memory a number of ways (e.g., flat file on disk, compressed sparse format, data frames, retrieved from SQL or NoSQL databases), our evaluation focused on computational performance rather than I/O. Our results show the relative strengths of each framework.

Graph algorithms, Benchmarking, shared-memory prog↗

DUNE Software and High Performance Computing

DUNE, like other HEP experiments, faces a challenge related to matching execution patterns of our production simulation and data processing software to the limitations imposed by modern high-performance computing facilities. In order to efficiently exploit these new architectures, particularly those with high CPU core counts and GPU accelerators, our existing software execution models require adaptation. In addition, the large size of individual units of raw data from the far detector modules pose an additional challenge somewhat unique to DUNE. Here we describe some of these problems and how we begin to solve them today with existing software frameworks and toolkits. We also describe ways we may leverage these existing software architectures to attack remaining problems going forward. This whitepaper is a contribution to the Computational Frontier of Snowmass21.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

GrainGNN: A dynamic graph neural network for predicting 3D grain microstructure

We propose GrainGNN, a surrogate model for the evolution of polycrystalline grain structure under rapid solidification conditions in metal additive manufacturing. High fidelity simulations of solidification microstructures are typically performed using multicomponent partial differential equations (PDEs) with moving interfaces. The inherent randomness of the PDE initial conditions (grain seeds) necessitates ensemble simulations to predict microstructure statistics, e.g., grain size, aspect ratio, and crystallographic orientation. Here, currently such ensemble simulations are prohibitively expensive and surrogates are necessary.In GrainGNN, we use a dynamic graph to represent interface motion and topological changes due to grain coarsening. We use a reduced representation of the microstructure using hand-crafted features; we combine pattern finding and altering graph algorithms with two neural networks, a classifier (for topological changes) and a regressor (for interface motion). Both networks have an encoder-decoder architecture; the encoder has a multi-layer transformer long-short-term-memory architecture; the decoder is a single layer perceptron.We evaluate GrainGNN by comparing it to high-fidelity phase field simulations for in-distribution and out-of-distribution grain configurations for solidification under laser power bed fusion conditions. GrainGNN results in 80%–90% pointwise accuracy; and nearly identical distributions of scalar quantities of interest (QoI) between phase field and GrainGNN simulations compared using Kolmogorov-Smirnov test. GrainGNN's inference speedup (PyTorch on single x86 CPU) over a high-fidelity phase field simulation (CUDA on a single NVIDIA A100 GPU) is 150×–2000× for 100-initial grain problem. Further, using GrainGNN, we model the formation of 11,600 grains in 220 seconds on a single CPU core.

36 MATERIALS SCIENCE↗