Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Architecture patterns”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Qualitative and quantitative analysis of neutron irradiation effects in SiC/SiC composites using X-ray computed tomography

Silicon carbide fiber-reinforced silicon carbide matrix (SiC/SiC) composites are candidate materials for cladding of light water reactor (LWR) fuels. Loss of fission product gas retention due to the formation of microcrack networks is considered a potential failure mechanism for SiC/SiC-cladded fuels. In this work, a variety of SiC/SiC composite tubes were irradiated with and without an LWR-relevant radial heat flux in the High Flux Isotope Reactor, followed by detailed characterization with X-ray computed tomography (XCT). This first set of XCT data for neutron-irradiated samples confirmed that the internal stresses arising from a combination of temperature gradients and irradiation-induced swelling act as the primary driver for cracking. Consequently, while the observed cracking patterns varied depending on the tube architectures, the sharp edges of relatively large pores were found to be the common stress concentrator. These findings are useful to help improve the design and manufacturing of SiC/SiC fuel claddings for reduced failure probability.

42 ENGINEERING↗

A New Semistructured Algebraic Multigrid Method

Multigrid methods are well suited to large massively parallel computer architectures because they are mathematically optimal and display good parallelization properties. Since current architecture trends are favoring regular compute patterns to achieve high performance, the ability to express structure has become much more important. The hypre software library provides high-performance multigrid preconditioners and solvers through conceptual interfaces, including a semistructured interface that describes matrices primarily in terms of stencils and logically structured grids. This paper presents a new semistructured algebraic multigrid (SSAMG) method built on this interface. The numerical convergence and performance of a CPU implementation of this method are evaluated for a set of semistructured problems. In conclusion, SSAMG achieves significantly better setup times than hypre’s unstructured AMG solvers and comparable convergence. In addition, the new method is capable of solving more complex problems than hypre’s structured solvers.

97 MATHEMATICS AND COMPUTING↗

Advanced Electrode Structures for Proton Exchange Membrane Fuel Cells: Current Status and Path Forward

Abstract Proton exchange membrane fuel cells (PEMFCs) have demonstrated their viability as a promising candidate for clean energy applications. However, performance of conventional PEMFC electrodes, especially the cathode electrode, suffers from low catalyst utilization and sluggish mass transport due to the randomly distributed components and tortuous transport pathways. Development of alternative architectures in which the electrode structure is controlled across a range of length scales provides a promising path toward overcoming these limitations. Here, we provide a comprehensive review of recent research and development of advanced electrode structures, organized by decreasing length-scale from the millimeter-scale to the nanometer-scale. Specifically, advanced electrode structures are categorized into five unique architectures for specific functions: (1) macro-patterned electrodes for enhanced macro-scale mass transport, (2) micro-patterned electrodes for enhanced micro-scale mass transport, (3) electrospun electrodes with fiber-based morphology for enhanced in-plane proton transport and through-plane O 2 transport, (4) enhanced-porosity electrodes for improved oxygen transport through selective inclusion of void space, and (5) catalyst film electrodes for elimination of carbon corrosion and ionomer poisoning. The PEMFC performance results achieved from each alternative electrode structure are presented and tabulated for comparison with conventional electrode architectures. Moreover, analysis of mechanisms by which new electrode structures can improve performance is presented and discussed. Finally, an overview of current limitations and future research needs is presented to guide the development of electrode structures for next generation PEMFCs. Graphical Abstract Development of improved electrode architectures with the control of structure on length scales ranging from millimeters to nanometers could enable a new generation of fuel cells with increased performance and reduced cost. This paper presents an in-depth review and critical analysis of recent developments and future outlook on the design of advanced electrode structures.

25 ENERGY STORAGE↗

Patterning of magneto-optical nanomaterials

Patterning of colloidal particles in precisely organized architectures has attracted intense research interest for decades. This is due to their potential applications in flexible electronics, magnetic and optical devices, sensors, biotechnology, communications, etc. However, creation of mesoscale assemblies at commercial scales have received less attention. The mesoscale systems reside between the micro- and macroscopic scales, with length dimensions from ≈ 100 µm to 5 mm. By leveraging decades of experimental and theoretical research in nanomaterial fields, we were able to precisely create and control the placement of nanoscale materials, allowing us to create mesoscale materials. We developed a versatile and automatic mesoscale patterning technology (via SEM-FIB and 3D printing) that provides precise and consistent control and special arrangement of functional nanomaterials. The versatility of the strategy is demonstrated by patterning nanoparticles with different dimensions, shapes and compositions, tethered with various functionalities and subjected to different external stimuli.

42 ENGINEERING↗

Directed self-assembly of block copolymers for high-precision patterning in the era of extreme ultraviolet lithography

Extreme ultraviolet (EUV) lithography enables unprecedented resolution in semiconductor patterning but faces critical challenges in developing resist materials that achieve high-precision at economically viable throughput. Directed self-assembly (DSA) of block copolymers (BCPs) offers a promising solution for pattern rectification by leveraging thermodynamically determined domain structures to decouple BCP pattern quality from the imperfect original lithographic pattern. This prospective presents an overview of recent progress on the EUV + DSA strategy, covering advances in BCP material design, processing, metrology, and pattern transfer. We highlight recent advances in high-χ BCPs with perpendicular orientation and domain spacings compatible with EUV dimensions, leveraging A-b-(B-r-C) architectures. We also discuss progress in chemical pre-pattern fabrication using both positive- and negative tone resists, along with processing strategies to minimize defects and roughness based on BCP thermodynamics and assembly kinetics. We further examine metrology platforms for characterizing the thermodynamics of BCP materials and quantifying the size and shape of BCP domains. Lastly, we review pattern transfer strategies for generating functional inorganic masks suitable for semiconductor manufacturing. Together, these advances highlight the potential of DSA to complement EUV lithography, offering a pathway to address critical challenges in achieving high-precision patterning for the semiconductor industry.

Lee, Kyunghyeon [Univ. of Chicago, IL (United Stat↗

Achieving performance portability in Gaussian basis set density functional theory on accelerator based architectures in NWChemEx

The numerical integration of the exchange–correlation (XC) potential is one of the primary computational bottlenecks in Gaussian basis set Kohn–Sham density functional theory (KS-DFT). To achieve optimal performance and accuracy, care must be taken in this numerical integration to preserve local sparsity as to allow for near linear weak scaling with system size. This leads to an integration scheme with several performance critical kernels which must be hand optimized for each architecture of interest. As the set of available accelerator hardware goes more diverse, a key challenge for developers of KS-DFT software is to maintain performance portability across a wide range of computational architectures. In this article, we examine a modular software design pattern which decouples the implementation details of performance critical kernels from the expression of high-level algorithmic workflows in a device-agnostic language such as C++; thus allowing for developers to target existing and emerging accelerator hardware within a single code base. We consider the efficacy of such a design pattern in the numerical integration of the XC potential by demonstrating its ability to achieve performance portability across a set of accelerator architectures which are representative of those on current and future U.S. Department of Energy Leadership Computing Facilities.

97 MATHEMATICS AND COMPUTING↗

Interfacial engineering via laser ablation for high-performing PEM water electrolysis

A rationalized interfacial design strategy was applied to tailor the porous transport layer (PTL)-catalyst layer (CL) contact and the PTL bulk-phase architecture. Particularly, at the PTL-CL interface, our results reveal that laser ablated sintered titanium power-based PTLs improve electrolyzer performance at both the H2NEW Consortium baseline catalyst loading of 0.4 mgIr cm -2 as well as at the ultra-low catalyst loading of 0.055 mgIr cm -2 . Under ultra-low catalyst loadings, the laser ablated PTL demonstrates maximum reduction of 230 mV compared to the commercial PTL at 4 A cm -2 , and reduces by 68 mV at 3.2 A cm -2 under H2NEW baseline loading. Laser ablation alters the titanium phase at the interface, so it forms more uniform structure like a microporous layer or a backing layer, leading to an increase in the surface area in contact with the catalyst layer while preventing the membrane from deforming into the PTL. Moreover, we reveal that bulk-phase architecture modification of the PTL by ablating patterned pores at the flow field-PTL interface improves mass transport without sacrificing contact at the CL-PTL interface. Overall, laser ablation of the PTL is an effective method to customize interfacial design to enhance proton exchange membrane electrolyzer performance.

08 HYDROGEN↗

Long-term compost amendment modulates wheat genotype differences in belowground carbon allocation, microbial rhizosphere recruitment and nitrogen acquisition

The implementation of soil health-promoting practices, such as cover cropping and compost application, has important implications for nutrient cycling and management in agroecosystems. At the same time, plant belowground carbon (C) allocation patterns can influence nutrient cycling and availability in soil through changes to the microbial community, but the effects may depend on the crop genotype and management practices in place. We evaluated belowground C allocation patterns using 13 C labeling and root architecture in two genotypes of winter wheat (Triticum aestivum) with different levels of exudation and belowground allocation strategies in soils with contrasting compost amendment legacy (108.7 Mg ha -1 every 2 years over 10 years vs. no compost). We also measured microbial community structure and function in the rhizosphere and quantified uptake of residue-derived N from 15 N-labelled cover crop residues. We found an interactive effect between soil management and genotype, where in the no-compost soil, the high-exudation genotype (Snowmass) increased exudation by over 4-fold, while the low-exudate genotype (Byrd) increased only 2-fold. While we did not observe genotype differences in rhizosphere enzyme activity or dissolved N pools, residue N uptake was 1.8 times greater for Snowmass in the compost-amended soil. There were more rhizosphere microbial taxa associated with the high-exudate genotype (Snowmass); nine bacterial and seven fungal families were indicative of Snowmass, versus one bacterial and four fungal families for Byrd. Our results suggest that the high-exudation strategy can influence the rhizosphere microbial community, and lead to greater short-term residue N uptake in high SOM soil. By directly linking root architecture, exudation, microbial communities, and N mineralization and uptake dynamics, this work demonstrates that plasticity in root C allocation is genotype-specific and influences microbial communities and nutrient cycling depending on the soil health context.

59 BASIC BIOLOGICAL SCIENCES↗

Self-assembled reconfigurable pump architectures via magnetic colloidal swarms

Self-assembled swarms of interactive active units, which are adaptive and dynamically reconfigurable to accommodate different functionalities, represent a promising platform for the development of next-generation robotics. Here, we utilize the emergent collective behavior of active magnetic colloids confined in quasi-two-dimensional arrays of overlapping wells to demonstrate the self-organization of a colloidal swarm into a dynamic pump architecture capable of controlled transport of passive cargo particles. This dynamic architecture provides a global unidirectional looping flow pattern along the entire length of the system. We show that the flow direction of the dynamic swarm-based pump can be externally controlled by a phase shift of a driving magnetic field energizing the swarm. The experimental observations are supported by computational modeling based on phenomenological coarse-grained particle dynamics coupled to shallow-water Navier-Stokes hydrodynamics. In conclusion, our findings demonstrate how the emergent collective behavior of a swarm can be orchestrated into a desired functionality by exploiting the interplay between activity and confinement potentials.

36 MATERIALS SCIENCE↗

Interpreting Write Performance of Supercomputer I/O Systems with Regression Models

This work seeks to advance the state of the art in HPC I/O performance analysis and interpretation. In particular, we demonstrate effective techniques to: (1) model output performance in the presence of I/O interference from production loads; (2) build features from write patterns and key parameters of the system architecture and configurations; (3) employ suitable machine learning algorithms to improve model accuracy. We train models with five popular regression algorithms and conduct experiments on two distinct production HPC platforms. We find that the lasso and random forest models predict output performance with high accuracy on both of the target systems. We also explore use of the models to guide adaptation in I/O middleware systems, and show potential for improvements of at least 15% from model-guided adaptation on 70% of samples, and improvements up to 10× on some samples for both of the target systems.

Xie, Bing↗

Time Series Foundation Models and Deep Learning Architectures for Earthquake Temporal and Spatial Nowcasting

Advancing the capabilities of earthquake nowcasting, the real-time forecasting of seismic activities, remains crucial for reducing casualties. This multifaceted challenge has recently gained attention within the deep learning domain, facilitated by the availability of extensive earthquake datasets. Despite significant advancements, the existing literature on earthquake nowcasting lacks comprehensive evaluations of pre-trained foundation models and modern deep learning architectures; each focuses on a different aspect of data, such as spatial relationships, temporal patterns, and multi-scale dependencies. This paper addresses the mentioned gap by analyzing different architectures and introducing two innovative approaches called Multi Foundation Quake and GNNCoder. We formulate earthquake nowcasting as a time series forecasting problem for the next 14 days within 0.1-degree spatial bins in Southern California. Earthquake time series are generated using the logarithm energy released by quakes, spanning 1986 to 2024. Our comprehensive evaluations demonstrate that our introduced models outperform other custom architectures by effectively capturing temporal-spatial relationships inherent in seismic data. The performance of existing foundation models varies significantly based on the pre-training datasets, emphasizing the need for careful dataset selection. However, we introduce a novel method, Multi Foundation Quake, that achieves the best overall performance by combining a bespoke pattern with Foundation model results handled as auxiliary streams.

97 MATHEMATICS AND COMPUTING↗

Understanding and Leveraging the I/O Patterns of Emerging Machine Learning Analytics

The scientific community is currently experiencing unprecedented amounts of data generated by cutting-edge science facilities. Soon facilities will be producing up to 1 PB/s which will force scientist to use more autonomous techniques to learn from the data. The adoption of machine learning methods, like deep learning techniques, in large-scale workflows comes with a shift in the workflow’s computational and I/O patterns. These changes often include iterative processes and model architecture searches, in which datasets are analyzed multiple times in different formats with different model configurations in order to find accurate, reliable and efficient learning models. This shift in behavior brings changes in I/O patterns at the application level as well at the system level. These changes also bring new challenges for the HPC I/O teams, since these patterns contain more complex I/O workloads. In this paper we discuss the I/O patterns experienced by emerging analytical codes that rely on machine learning algorithms and highlight the challenges in designing efficient I/O transfers for such workflows. We comment on how to leverage the data access patterns in order to fetch in a more efficient way the required input data in the format and order given by the needs of the application and how to optimize the data path between collaborative processes. We will motivate our work and show performance gains with a study case of medical applications.

Gainaru, Ana↗

HARMONY: Large-Scale Architecture Search for Efficient Hybrid Language Models

As large language models scale to trillions of parameters, their computational and memory requirements present critical challenges for efficient training and deployment. While Mixture of Experts (MoE) architectures enable efficient scaling through sparse parameter activation, and state-space models like Mamba offer linear-time complexity, principled methods for combining these paradigms remain undeveloped. We introduce HARMONY (Hybrid Architecture Research for Mamba, Optimized with Neural efficiencY), a multi-objective evolutionary neural architecture search framework for discovering efficient hybrid language models that integrate Transformer attention mechanisms, Mixture-of-Experts routing, and Mamba state-space components. Through large-scale distributed search using 16,384 MI250X GPUs on the Frontier supercomputer, HARMONY explores a comprehensive design space encompassing six attention variants (MHA, MQA, GQA, MLA, SWA, and Mamba-2), variable MoE configurations with both routed and shared experts, and extensive Mamba hyperparameters. Our framework discovers heterogeneous architectures that balance training performance with computational efficiency through multi-objective optimization incorporating latency penalties and fitness-based selection. Analysis of discovered architectures reveals that optimal hybrid designs favor heterogeneous component mixing rather than homogeneous patterns, with Mamba-2 and Multi-Head Latent Attention (MLA) emerging as preferred mechanisms. Discovered architectures demonstrate superior training efficiency: our best configuration achieves a final perplexity of 1.0874 with 2.38B parameters while processing 4,320 tokens/second, outperforming significantly larger manually designed models. Full-scale evaluation shows HARMONY's top architectures achieve better loss trajectories than equivalently-sized models using state-of-the-art configurations including Mixtral, Jamba, and Samba. Additionally, we demonstrate 91% weak scaling efficiency when training discovered 36B-parameter models across 1,024 GPUs. HARMONY is released as an open framework with comprehensive tools for building and training hybrid models using expert-data-pipeline parallelism, democratizing access to automated architecture design for next-generation language models.

Herron, Emily [ORNL] (ORCID:0000000273008172)↗

Polymeric multimaterials by photochemical patterning of crystallinity

An organized combination of stiff and elastic domains within a single material can synergistically tailor bulk mechanical properties. However, synthetic methods to achieve such sophisticated architectures remain elusive. We report a rapid, facile, and environmentally benign method to pattern strong and stiff semicrystalline phases within soft and elastic matrices using stereo-controlled ring-opening metathesis polymerization of an industrial monomer, cis -cyclooctene. Dual polymerization catalysis dictates polyolefin backbone chemistry, which enables patterning of compositionally uniform materials with seamless stiff and elastic interfaces. Visible light–induced activation of a metathesis catalyst results in the formation of semicrystalline trans polyoctenamer rubber, outcompeting the formation of cis polyoctenamer rubber, which occurs at room temperature. This bottom-up approach provides a method for manufacturing polymeric materials with promising applications in soft optoelectronics and robotics.

Science & Technology - Other Topics↗

Interpretable Models for Workflow Differentiation in High-Performance Scientific Networks

Scientific workflows in high-performance networks spawn hundreds of interdependent flows that must be managed collectively—yet existing network classifiers treat each flow in isolation, leading to fragmented QoS decisions and missed interflow patterns. We present a novel traffic classification solution that operates at the workflow level, distinguishing entire filetransfer operations from streaming analytics by capturing how concurrent flows interact and burst together. We introduce a workflow identification window (WIW) that ingests raw packet headers from parallel flows into unified tensors, preserving the spatial-temporal patterns that differentiate scientific workflows. This approach achieves 98.7% accuracy using CNN, LSTM, and hybrid architectures, while maintaining 84% accuracy on production traffic collected a week later—demonstrating robustness to temporal drift. By integrating SHAP and GradCAM explainability, we reveal that early-packet timing patterns and cross-flow correlations drive classification decisions, providing operators with interpretable insights. Our system enables coherent workflow-level QoS enforcement and dynamic bandwidth allocation in scientific networks, eliminating manual per-flow configuration while maintaining classification latency at millisecond level.

Giannakou, Anna [LBL, Berkeley]↗

Hardware-Based Emulator with Deep Learning Model for Building Energy Control and Prediction Based on Occupancy Sensors’ Data

Heating, ventilation, and air conditioning (HVAC) is the largest source of residential energy consumption. Occupancy sensors’ data can be used for HVAC control since it indicates the number of people in the building. HVAC and sensors form a typical cyber-physical system (CPS). In this paper, we aim to build a hardware-based emulation platform to study the occupancy data’s features, which can be further extracted by using machine learning models. In particular, we propose two hardware-based emulators to investigate the use of wired/wireless communication interfaces for occupancy sensor-based building CPS control, and the use of deep learning to predict the building energy consumption with the sensor data. We hypothesize is that the building energy consumption may be predicted by using the occupancy data collected by the sensors, and question what type of prediction model should be used to accurately predict the energy load. Another hypothesis is that an in-lab hardware/software platform could be built to emulate the occupancy sensing process. The machine learning algorithms can then be used to analyze the energy load based on the sensing data. To test the emulator, the occupancy data from the sensors is used to predict energy consumption. The synchronization scheme between sensors and the HVAC server will be discussed. We have built two hardware/software emulation platforms to investigate the sensor/HVAC integration strategies, and used an enhanced deep learning model—which has sequence-to-sequence long short-term memory (Seq2Seq LSTM)—with an attention model to predict the building energy consumption with the preservation of the intrinsic patterns. Because the long-range temporal dependencies are captured, the Seq2Seq models may provide a higher accuracy by using LSTM architectures with encoder and decoder. Meanwhile, LSTMs can capture the temporal and spatial patterns of time series data. The attention model can highlight the most relevant input information in the energy prediction by allocating the attention weights. The communication overhead between the sensors and the HVAC control server can also be alleviated via the attention mechanism, which can automatically ignore the irrelevant information and amplify the relevant information during CNN training. Our experiments and performance analysis show that, compared with the traditional LSTM neural network, the performance of the proposed method has a 30% higher prediction accuracy.

Ye, Zhijing↗

Architecture and performance of Perlmutter's 35 PB ClusterStor E1000 all-flash file system

NERSC's newest system, Perlmutter, features a 35 PB all-flash Lustre file system built on HPE Cray ClusterStor E1000. Here, we present its architecture, early performance figures, and performance considerations unique to this architecture. We demonstrate the performance of E1000 OSSes through low-level Lustre tests that achieve over 90% of the theoretical bandwidth of the SSDs at the OST and LNet levels. We also show end-to-end performance for both traditional dimensions of I/O performance (peak bulk-synchronous bandwidth) and nonoptimal workloads endemic to production computing (small, incoherent I/Os at random offsets) and compare them to NERSC's previous system, Cori, to illustrate that Perlmutter achieves the performance of a burst buffer and the resilience of a scratch file system. Finally, we discuss performance considerations unique to all-flash Lustre and present ways in which users and HPC facilities can adjust their I/O patterns and operations to make optimal use of such architectures.

97 MATHEMATICS AND COMPUTING↗

Accurate and Data‐Efficient Micro X‐ray Diffraction Phase Identification Using Multitask Learning: Application to Hydrothermal Fluids

Traditional analysis of highly distorted micro X‐ray diffraction (μ‐XRD) patterns from hydrothermal fluid environments is a time‐consuming process, often requiring substantial data preprocessing and labeled experimental data. Herein, the potential of deep learning with a multitask learning (MTL) architecture to overcome these limitations is demonstrated. MTL models are trained to identify phase information in μ‐XRD patterns, minimizing the need for labeled experimental data and masking preprocessing steps. Notably, MTL models show superior accuracy compared to binary classification convolutional neural networks. Additionally, introducing a tailored cross‐entropy loss function improves MTL model performance. Most significantly, MTL models tuned to analyze raw and unmasked XRD patterns achieve close performance to models analyzing preprocessed data, with minimal accuracy differences. This work indicates that advanced deep learning architectures like MTL can automate arduous data handling tasks, streamline the analysis of distorted XRD patterns, and reduce the reliance on labor‐intensive experimental datasets.

97 MATHEMATICS AND COMPUTING↗