Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Architecture patterns”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Machine learning for domain transfer between simulated and experimental 2D X-ray diffraction patterns using generative adversarial networks

X-ray diffraction (XRD) is a well-established technique for analyzing materials at an atomic level. Dynamic compression experiments (DCE), in which materials are subject to extreme pressures, can provide fundamental understanding to pressure-induced phase transitions and compression of the crystal lattice. The analysis of XRD patterns from highly compressed samples is non-trivial given the sparsity of data, high experimental costs, and the fact that the data is often marred with X-ray background and other artifacts. While accurate computational frameworks exist, they solve the forward problem—from structures and orientations to XRD patterns. Solving the inverse problem for 2D experimental diffraction patterns is currently a complex manual process of matching and comparing experimentally observed patterns to computationally generated ones. Machine learning is a promising tool for automating the matching process but often requires data-intensive architectures. Here, in this study, we use a CycleGAN to translate the domain of limited experimental data to a domain in which there is readily available simulated data. This domain shift allows data-intensive machine learning models that have only been trained on simulated XRD patterns to be used in the analysis of experiments.

Brozak, Samantha Jean [Sandia National Laboratorie↗

NeuroCoreX: An Open-Source FPGA-Based Spiking Neural Network Emulator with On-Chip Learning

Spiking Neural Networks (SNNs) are computational models inspired by the event-driven communication and connectivity patterns of biological neural circuits. They enable high energy efficiency and natural support for diverse architectures ranging from layered networks to small-world and graphstructured topologies. In this work, we introduce NeuroCoreX, an open-source, FPGA-based spiking neural network emulator that provides real-time, on-chip learning and flexible network organization. NeuroCoreX supports both feedforward sensory inputs streamed directly from sensors or PCs via UART and recurrent on-chip connectivity, enabling simultaneous processing and learning from external stimuli and internal network dynamics-capabilities rarely available in existing FPGA SNN platforms. The system implements a Leaky Integrate-and-Fire (LIF) neuron model with current-based synapses and supports pair-based STDP learning on both feedforward and recurrent synapses. A lightweight Python interface enables interactive configuration, live monitoring, weight read-back, and experiment control. Importantly, NeuroCoreX is tightly integrated with the SuperNeuroMAT simulator, allowing SNN models to be transferred seamlessly from software to hardware for hardware-in-the-loop development. By combining real-time plasticity, flexible connectivity, and an open-source VHDL implementation, NeuroCoreX provides an extensible and accessible platform for neuromorphic research, algorithm-hardware co-design, and energy-efficient edge intelligence.

Gautam, Ashish [ORNL]↗

Experimental Investigation of Buoyant Flow in Realistic Bedforms With Heterogeneous Wettability

Submeter-scale geologic heterogeneity greatly affects CO 2 plume migration and retention. In this work, we present meter-scale laboratory experiments that can capture the impact of realistic submeter-scale geologic heterogeneity on multiphase flow and trapping. We produce realistic sedimentary formations consisting of ripple deposits with varying grain size contrast and wettability in a meter-scale slab chamber. Then, we conduct multiphase flow experiments with analog fluids through these structures and measure the saturation patterns, capillary heterogeneity trapping (CHT), and overall trapping performance. When we alter the ripple bedform architecture, variations in trapped saturation and CHT (10–20%) increment are exhibited. Similar growth in trapping performance is also observed when grain size contrast increases. Finally, wettability changes (water- to oil-wet) can increase nonwetting saturation and CHT up to 5% and 10–20%, respectively. These results emphasize the importance of correctly characterizing the impact of small-scale heterogeneities and wettability changes. We believe this is the first time that multiphase flow experiments were conducted in meter-scale domains with realistic ripple bedforms and heterogeneous wettability to investigate plume migration and trapping.

58 GEOSCIENCES↗

Comparing LLC-Memory Traffic between CPU and GPU Architectures

The cache hierarchy in modern CPUs and GPUs is becoming increasingly complex, which makes understanding the handshake between the memory access patterns and the cache hierarchy difficult. Moreover, the details of different cache policies are not publicly available. Therefore, the research community relies on observation to understand the relationship between memory access patterns and cache hierarchy. Our previous studies delved into the different microarchitectures of Intel CPUs. In this study, GPUs from NVIDIA and AMD are considered. Even though the execution models in CPUs and GPUs are distinct, this study attempts to correlate the behavior of the cache hierarchy of CPUs and GPUs. Using the knowledge gathered from studying Intel CPUs, the similarities and dissimilarities between CPUs and GPUs are identified. Through model evaluation, this study provides a proof of concept that traffic between last-level cache and memory can be predicted for sequential streaming and strided access patterns on GPUs.

Monil, M. A. H.↗

Area-selective deposition of germanium on patterned graphene/monolayer molybdenum disulfide stacks via dipole engineering

Heterogeneous integration of two-dimensional materials and the conventional semiconductor has opened opportunities for next-generation semiconductor devices and their processing. Heterogeneous integration has been studied for economical manufacturing by substrate recycling and novel functionalities by a combination of incommensurate materials. However, utilizing the integration requires controlling locations of the integrated architectures. Here, we show area-selective deposition (ASD) of germanium on the graphene/MoS 2 stack. Ge nucleation precisely occurred on the surfaces of the patterned graphene/MoS 2 stack via dipole engineering. In this study, the growth temperature of ASD of Ge was significantly lower than that based on precursor desorption on SiO 2 . The first-principles calculations revealed that Ge deposited by ASD on the graphene/MoS 2 stack was not affected by charge transfer. This work provides a viable way to utilize atomically thin materials for next-generation semiconductor devices, which can be applicable for “Beyond Moore” and “More Moore” approaches.

2D materials↗

Discovering and Demonstrating a Novel High-Performing 2D-Patterned Electrode for Proton-Exchange Membrane Water Electrolysis Devices

Proton-exchange membrane water electrolysis (PEMWE) produces hydrogen with high efficiency and purity but uses high-loading platinum-group metal (PGM) catalysts. Such concerns call for the development of novel electrode architectures to improve catalyst utilization and mass activity, thus promoting PEMWE cost competitiveness for large-scale implementation. In this study, we demonstrated, for the first time, a novel two-dimensional (2D)-patterned electrode with edge effects to address these challenges. The edge effect was induced by membrane properties, potential distribution, and counter electrode coverage and could be optimized by tuning the catalyst layer dimensions. To achieve identical PEMWE performance, the optimal pattern saved the 21% anode PGM catalyst compared with the conventional catalyst fully covered electrode. The PGM catalyst could be further reduced by 61% to boost mass activity with no significant performance loss. The results also indicated that the electrode uniformity in PEMWE cells might not be as critical as that in PEM fuel cells. Finally, the novel 2D-patterned electrode could effectively reduce PGM catalyst loading, accelerating affordable and large-scale production of hydrogen and other value-added chemicals via electrolysis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Structure–Function Relationships in Sequence-Controlled Copolymers for Rare Earth Element Chelation

The ability to tune material function through primary sequence is a defining feature of biological macromolecules, allowing precise control over structure and target interactions in complex aqueous environments. However, translating sequence–structure–function relationships to synthetic macromolecules is challenging due to their dispersity in sequence, conformation, and composition. Here, we report systematic studies of amphiphilic polymer chelators designed to probe how composition and patterning influence binding affinity and selectivity for rare earth elements (REEs), a series of technologically relevant metals with challenging separation profiles. A library of copolymers varying hydrophobic monomer composition and patterning was synthesized via reversible addition–fragmentation chain transfer (RAFT) polymerization, spanning statistical, gradient, and block architectures. REE binding was quantified using a high-throughput colorimetric assay, and reconstruction of polymer ensembles using kinetic stochastic simulations enabled quantitative comparisons of sequence heterogeneity, linking local monomer colocalization to emergent REE binding. Further, we investigated the role of different hydrophobic comonomers in tuning metal coordination, with binding trends linked to structural features that influence binding site desolvation. Complementary dynamic light scattering (DLS) and small-angle X-ray scattering (SAXS) measurements showed that both polymer and monomer architecture modulate metal-induced conformational changes, and that multichain assembly behavior emerges beyond critical hydrophobic thresholds. Sequence control also altered REE selectivity, with nonmonotonic differences observed across compositionally identical polymers with different sequence architectures. Together, these findings establish design principles that connect polymer sequence and structure to binding performance, guiding the design of macromolecular chelators with enhanced affinity and selectivity for applications in separations, sensing, and catalysis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Crowdsourcing the Frontier: Advancing Hybrid Physics‐ML Climate Simulation via a $\$$50,000 Kaggle Competition

Subgrid machine-learning (machine learning [ML]) parameterizations have the potential to introduce a new generation of climate models that incorporate the effects of higher-resolution physics without incurring the prohibitive computational cost associated with more explicit physics-based simulations. However, important issues, ranging from online instability to inconsistent online performance, have limited their operational use for long-term climate projections. To more rapidly drive progress in solving these issues, domain scientists and ML researchers opened up the offline aspect of this problem to the broader ML and data science community with the release of ClimSim, a NeurIPS Data sets and Benchmarks publication, and an associated Kaggle competition. This paper reports on the downstream results of the Kaggle competition by coupling emulators inspired by the winning teams' architectures to an interactive climate model (including full cloud microphysics, a regime historically prone to online instability) and systematically evaluating their online performance. Our results demonstrate that online stability in the low-resolution real-geography setting is reproducible across multiple diverse architectures, which we consider a key milestone. All tested architectures exhibit strikingly similar offline and online biases, though their responses to architecture-agnostic design choices (e.g., expanding the list of input variables) can differ significantly. Multiple Kaggle-inspired architectures achieve state-of-the-art results on certain metrics such as zonal mean bias patterns and global Root Mean Squared Error, indicating that crowdsourcing the essence of the offline problem is one path to improving online performance in hybrid physics-AI climate simulation.

Environmental sciences↗

Sparse Linear Solvers for Large-scale Electromagnetic Transient Simulations

Linear solvers form the basis for electromagnetic transient (EMT) simulations. There is a need to speed up EMT simulations as larger regions are analyzed using EMT simulations. For the same, the performance of linear solvers plays an important role. Exploiting the sparsity of the matrices generated in EMT simulations could assist with speed-up. Scalability is also crucial as power grids expand, demanding solutions capable of accommodating the increasing system size. Recent studies from the North American Electric Reliability Corporation (NERC) increasingly emphasize that EMT simulation models of the power grid will grow larger with the inclusion of power electronics components. Parallelisms in sparsity patterns exploit modern central processing units (CPUs), multi-core CPUs, and graphics processing units (GPUs) architectures in sparse solver designs. Therefore, this paper explores publicly available existing linear solvers and investigates their efficiency in large-scale power grid simulations. A large-scale power grid is developed by increasing the size of the IEEE 39 bus test system to up to 39000 bus systems.

Hsu, Kuan-Chieh↗

Self-Forming Thin Interphases and Electrodes Enabling 3-D Structured High Energy Density Batteries

An electrolytically in-situ formed fluoride/lithium based battery has been developed to offer a pathway to scalable reconfigurable solid state batteries of high energy density. Research into the development of novel in-situ formed chemistries encompassing the negative and positive reactive current collectors, and the bi-ion glass conductor along with electrode structure was accomplished with a focused attention on transport. The solid state in-situ batteries were fabricated with a maskless scalable patterning technique to offer a pathway to high throughput, low material loss and fabrication of complex architectures. Such development and integration enabled to achieve the 12 V bipolar batteries at > 1000 Wh/L energy density based on electrode pairs and current collectors.

25 ENERGY STORAGE↗

A Range and Performance Optimized Version of the Computer-Aided Speckle Interferometry Algorithm for Real-Time Displacement-Strain Field Monitoring

Abstract This work presents an optimized implementation of the Computer-Aided Speckle Interferometry algorithm which enables full-field determination of displacements and strains on commodity Graphics Processing Units at high resolution and frame rates. By combining careful control of the average speckle size in a laser speckle pattern with a simple sampling rate conversion scheme, a compact representation of the optical speckle is achieved. This allows for optimal use of Graphics Processing Unit architecture with robust range extension. The optimal mapping of the Computer-Aided Speckle Interferometry algorithm to Graphics Processing Unit architecture is shown in detail, and a straightforward method for disambiguating large displacements is illustrated. Lastly, this paper demonstrates a two-step subimage-tapering modification to the original algorithm that enables robust range enhancement while maintaining resolution. Results from numerical simulations on synthetic speckle patterns are shown, and runtime performance metrics are provided, with performance ranging up to 60 frames per second in some cases. The method is suitable for interactive experimental mechanics research, process and testing or any application where real-time high-resolution displacement-strain monitoring is needed. A .NET Framework class library enabling the incorporation of the algorithm into 3rd -party applications is available for download.

42 ENGINEERING↗

Towards Generic Parallel Programming in Computer Science Education with Kokkos

Parallel patterns, views, and spaces are promising abstractions to capture the programmer's intent as well as the contextual information that can be used by an underlying runtime to efficiently map software to parallel hardware. These abstractions can be valuable in cases where an algorithm must accommodate requirements of code and performance portability across hardware architectures and vendor programming models. Kokkos is a parallel programming model for host- and accelerator architectures that relies on these abstractions and targets these requirements. It consists of a pure C++ interface, a specification, and a programming library. The programming library exposes patterns and types and maps them to an underlying abstract machine model. The abstract machine model offers a generic view of parallel hardware. While Kokkos is gaining popularity in large-scale HPC applications at some DOE laboratories, we believe that the implemented concepts are of interest to a broader audience including academia as they may contribute to a generic, vendor, and architecture-independent education of parallel programming. In this work, we give an insight into the design considerations of this programming model and list important abstractions. Further, we document best practices obtained from giving virtual classes on Kokkos and give pointers to resources that the reader may consider valuable for a lecture on generic parallel programming for students with preexisting knowledge on this matter.

Ciesko, Jan↗

Multielectrode electrochemical cell for in situ structural characterization of amorphous thin-film catalysts using high-energy X-ray scattering

A multielectrode-based electrochemical cell allows the structural characterization of an amorphous thin-film water oxidation catalyst under various electrochemical potentials using high-energy X-ray scattering and atomic pair distribution function (PDF) techniques. A multielectrode with five electrodes provides a sufficiently low background signal to enable high-energy X-ray scattering (HEXS) measurements and amplifies the extremely low HEXS signals from samples for high-resolution PDF analysis of in situ data from thin-film catalysts. Glassy carbon (GC) creates a relatively low intensity HEXS pattern and is used as a working electrode. Instead of a three-dimensional (3D) porous electrode architecture, the flat geometry of the electrode enables various deposition techniques to be used for the preparation of a highly conductive metal oxide layer. PDF analysis demonstrates high spatial resolution for a 230 nm thick amorphous iridium oxide film deposited on two roughened 60 µm thick GC electrodes. The PDF analysis resolves the domain size and distinguishes changes in fine structure which are directly correlated with the structure and function of the catalysts. In conclusion, the results bring the opportunity to analyze the structure of nanometre-scale amorphous thin-film catalysts in an electrolyte-compatible and compact 3D-printed electrochemical cell in a three-electrode configuration.

36 MATERIALS SCIENCE↗

Enabling Efficient Sparse Computations using Linear Algebra Aware Compilers

This project developed the LAPIS compiler framework, built on the Multilevel Intermediate Representation (MLIR), to optimize sparse linear algebra operations and support performance portability across diverse architectures. The main innovation of LAPIS is the Kokkos dialect, which allows for lowering codes from a high productivity language to different architectures in an elegant way. The dialect also allows the conversion of lower-level MLIR code to C++ Kokkos code, facilitating the integration of scientific machine learning (SciML) models into applications. To extend LAPIS for distributed memory architectures, a new partition dialect was created to manage the distribution of sparse tensors and express communication patterns for sparse linear algebra operations. This dialect also supports the distributed execution of operators and includes algorithmic optimizations to minimize communication to improve performance. The project also demonstrates that MLIR can enable effective linear algebra-level optimizations, improving performance on different GPUs for both sparse and dense linear algebra kernels. Key applications of LAPIS include sparse linear algebra and graph kernels, TenSQL, a relational database management solution built on GraphBLAS, and the development of subgraph isomorphism and monomorphism kernels, showcasing performance portability. In summary, the LAPIS framework supports productivity, performance, portability, and distributed memory execution, while also enabling linear algebra-level optimizations that are challenging in traditional programming languages, with successful applications ranging from simple sparse linear algebra to complex graph kernels.

97 MATHEMATICS AND COMPUTING↗

An MLCommons Scientific Benchmarks Ontology

Scientific machine learning research spans diverse domains and data modalities, yet existing benchmark efforts remain siloed and lack standardization. This makes novel and transformative applications of machine learning to critical scientific use-cases more fragmented and less clear in pathways to impact. This paper introduces an ontology for scientific benchmarking developed through a unified, community-driven effort that extends the MLCommons ecosystem to cover physics, chemistry, materials science, biology, climate science, and more. Building on prior initiatives such as XAI-BENCH, FastML Science Benchmarks, PDEBench, and the SciMLBench framework, our effort consolidates a large set of disparate benchmarks and frameworks into a single taxonomy of scientific, application, and system-level benchmarks. New benchmarks can be added through an open submission workflow coordinated by the MLCommons Science Working Group and evaluated against a six-category rating rubric that promotes and identifies high-quality benchmarks, enabling stakeholders to select benchmarks that meet their specific needs. The architecture is extensible, supporting future scientific and AI/ML motifs, and we discuss methods for identifying emerging computing patterns for unique scientific workloads. The MLCommons Science Benchmarks Ontology provides a standardized, scalable foundation for reproducible, cross-domain benchmarking in scientific machine learning. A companion webpage for this work has also been developed as the effort evolves: https://mlcommons-science.github.io/benchmark/

Hawks, Ben [Fermilab] (ORCID:0000000157000288)↗

Anticipating Technical Expertise and Capability Evolution in Research Communities Using Dynamic Graph Transformers

The ability to anticipate global technical expertise and capability evolution trends is essential for national and global security, especially in safety-critical domains such as nuclear nonproliferation (NN) and rapidly emerging fields like artificial intelligence (AI). Here, in this work, we extend traditional statistical relational learning approaches (e.g., link prediction in collaboration networks) and formulate a problem of anticipating technical expertise and capability evolution using dynamic heterogeneous graph representations. We develop novel capabilities to forecast collaboration patterns, authorship behavior, and technical capability evolution at different granularities (e.g., scientist and institution levels) in two distinct research fields. We implement a dynamic graph transformer (DGT) neural architecture, which pushes the state-of-the-art graph neural network models by: 1) forecasting heterogeneous (rather than homogeneous) nodes and edges; and 2) relying on both discrete- and continuous-time inputs. We demonstrate that our DGT models predict collaboration, partnership, and expertise patterns with 0.26, 0.73, and 0.53 mean reciprocal rank values for AI and 0.48, 0.93, and 0.22 for NN domains. DGT model performance exceeds the best-performing static graph baseline models by 30%–80% across AI and NN domains. Our findings demonstrate that DGT models boost inductive task performance when previously unseen nodes appear in the test data for the domains with emerging collaboration patterns (e.g., AI). Specifically, models accurately predict which established scientists will collaborate with early career scientists and vice versa in the AI domain.

97 MATHEMATICS AND COMPUTING↗

DUNE Database Development

The DUNE experiment will produce vast amounts of metadata, which describe the data coming from the read-out of the primary DUNE detectors. Various databases will make up the overall DB architecture for this metadata. ProtoDUNE at CERN is the largest existing prototype for DUNE and serves as a testing ground for - among other things - possible database solutions for DUNE. The subset of all metadata that is accessed during offline data reconstruction and analysis is referred to as ‘conditions data’ and it is stored in a dedicated database. As offline data reconstruction and analysis will be deployed on HTC and HPC resources, conditions data is expected to be accessed at very high rates. It is therefore crucial to store it in a granularity that matches the expected access patterns allowing for extensive caching. This requires a good understanding of the sources and use cases of conditions data. This contribution will briefly summarize the database architecture deployed at ProtoDUNE and explain the various sources of conditions data. We will present how the conditions data is retrieved and streamed from the databases and how it is handled to match expected access patterns.

Vizcaya Hernandez, Ana Paula↗