Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “bottleneck structure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Interplay between electron localization, magnetic order, and Jahn-Teller distortion dictates LiMnO2 phase stability

The development of manganese (Mn)-rich cathodes for Li-ion batteries promises to alleviate potential supply chain bottlenecks in battery manufacturing. Fundamental challenges in Mn-rich cathodes arise from phenomena such as structural changes due to cooperative Jahn-Teller (JT) distortions of in octahedral environments, Mn migration, and phase transformations to spinel-like order, all of which affect the electrochemical performance. These physically complex phenomena motivate an re-examination of the Li-Mn-O rock-salt space, with a focus on the thermodynamics of the prototypical, polymorphs. It is found that the generalized gradient approximation (GGA-PBEsol) and meta-GGA ( ) density functionals with empirically fitted on-site Hubbard corrections yield spurious stable phases for , such as predicting a phase with -like order ( ) to be the ground state instead of the orthorhombic (Pmmn) phase, which is the experimentally known ground state. Accounting for antiferromagnetic order in each structure is shown to have a substantial effect on the total energies and resulting phase stability. By using hybrid-GGA (HSE06) and GGA with self-consistent Hubbard parameters (on-site and inter-site ) calculated from linear response theory, the experimentally observed phase stability trends are recovered. The calculated on-site between Mn- states in the experimentally observed orthorhombic, layered, and spinel phases are significantly smaller than in and disordered layered structures, by within GGA. The smaller values of are shown to be correlated with a collinear ordering of JT distortions, in which all orbitals are oriented in the same direction. This cooperative JT effect can lead to greater electron delocalization from Mn along the states due to increased Mn-O covalency, which contributes to the greater electronic stability compared to the phases with noncollinear JT arrangements. The structures with collinear ordering of JT distortions also generate greater vibrational entropy, which helps stabilize these phases at high temperature. These phases are shown to be strongly insulating with large calculated band gaps , which are computed using HSE06 and .

Kam, Ronald L↗

Scalable Incremental Checkpointing using GPU-Accelerated De-Duplication

Writing large amounts of data concurrently to stable storage is a typical I/O pattern of many HPC workflows. This pattern introduces high I/O overheads and results in increased storage space utilization especially for workflows that need to capture the evolution of data structures with high frequency as checkpoints. In this context, many applications, such as graph pattern matching, perform sparse updates to large data structures between checkpoints. For these applications, incremental checkpointing techniques that save only the differences from one checkpoint to another can dramatically reduce the checkpoint sizes, I/O bottlenecks, and storage space utilization. However, such techniques are not without challenges: it is non-trivial to transparently determine what data has changed since a previous checkpoint and assemble the differences in a compact fashion that does not result in excessive metadata. State-of-art data reduction techniques (e.g., compression and de-duplication) have significant limitations when applied to modern HPC applications that leverage GPUs: slow at detecting the differences, generate a large amount of metadata to keep track of the differences, and ignore crucial spatiotemporal checkpoint data redundancy. This paper addresses these challenges by proposing a Merkle tree-based incremental checkpointing method to exploit GPUs' high memory bandwidth and massive parallelism. Experimental results at scale show a significant reduction of the I/O overhead and space utilization of checkpointing compared with state-of-the-art incremental checkpointing and compression techniques.

Tan, Nigel↗

Magneto-optical study of Nb thin films for superconducting qubits

Abstract Among the recognized sources of decoherence in superconducting qubits, the spatial inhomogeneity of the superconducting state and the possible presence of magnetic-flux vortices remain comparatively underexplored. Niobium is commonly used as a structural material in transmon qubits that host Josephson junctions, and excess dissipation anywhere in the transmon can become a bottleneck that limits overall quantum performance. The metal/substrate interfacial layer may simultaneously host pair-breaking loss channels (e.g. two-level systems) and control thermal transport, thereby affecting dissipation and temperature stability. Here, we use quantitative magneto-optical imaging of the magnetic-flux distribution to characterize the homogeneity of the superconducting state and the critical current density, j c , in niobium films fabricated under different sputtering conditions. The imaging reveals distinct flux-penetration regimes, ranging from a nearly ideal Bean critical state to strongly nonuniform thermo-magnetic dendritic avalanches. By fitting the measured magnetic-induction profiles, we extract j c and try to correlate it with film physical properties and with measured qubit internal quality factors. Our results indicate that the Nb/Si interlayer can be a significant contributor to decoherence and should be considered an important factor that must be optimized.

Datta, Amlan [Ames National Laboratory; Iowa State↗

Multiorbital Quantum Impurity Solver for General Interactions and Hybridizations

Here we present a numerically exact inchworm Monte Carlo method for equilibrium multiorbital quantum impurity problems with general interactions and hybridizations. We show that the method, originally developed to overcome the dynamical sign problem in certain real-time propagation problems, can also overcome the sign problem as a function of temperature for equilibrium quantum impurity models. This is shown in several cases where the current method of choice, the continuous-time hybridization expansion, fails due to the sign problem. Our method therefore enables simulations of impurity problems as they appear in embedding theories without further approximations, such as the truncation of the hybridization or interaction structure or a discretization of the impurity bath with a set of discrete energy levels, and eliminates a crucial bottleneck in the simulation of ab initio embedding problems.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Magneto-optical study of Nb thin films for superconducting qubits

Among the recognized sources of decoherence in superconducting qubits, the spatial inhomogeneity of the superconducting state and the possible presence of magnetic-flux vortices remain comparatively underexplored. Niobium is commonly used as a structural material in transmon qubits that host Josephson junctions, and excess dissipation anywhere in the transmon can become a bottleneck that limits overall quantum performance. The metal/substrate interfacial layer may simultaneously host pair-breaking loss channels (e.g., two-level systems, TLS) and control thermal transport, thereby affecting dissipation and temperature stability. Here, we use quantitative magneto-optical imaging of the magnetic-flux distribution to characterize the homogeneity of the superconducting state and the critical current density, $j_{c}$, in niobium films fabricated under different sputtering conditions. The imaging reveals distinct flux-penetration regimes, ranging from a nearly ideal Bean critical state to strongly nonuniform thermo-magnetic dendritic avalanches. By fitting the measured magnetic-induction profiles, we extract $j_{c}$ and correlate it with film physical properties and with measured qubit internal quality factors. Our results indicate that the Nb/Si interlayer can be a significant contributor to decoherence and should be considered an important factor that must be optimized.

Datta, Amlan [Ames Lab; Iowa State U.]↗

Fast correlation function calculator: A high-performance pair-counting toolkit

A novel high-performance exact pair-counting toolkit called fast correlation function calculator (FCFC) is presented. With the rapid growth of modern cosmological datasets, the evaluation of correlation functions with observational and simulation catalogues has become a challenge. High-efficiency pair-counting codes are thus in great demand. We introduce different data structures and algorithms that can be used for pair-counting problems, and perform comprehensive benchmarks to identify the most efficient algorithms for real-world cosmological applications. We then describe the three levels of parallelisms used by FCFC, SIMD, OpenMP, and MPI, and run extensive tests to investigate the scalabilities. Finally, we compare the efficiency of FCFC with alternative pair-counting codes. The data structures and histogram update algorithms implemented in FCFC are shown to outperform alternative methods. FCFC does not benefit greatly from SIMD because the bottleneck of our histogram update algorithm is mainly cache latency. Nevertheless, the efficiency of FCFC scales well with the numbers of OpenMP threads and MPI processes, even though speedups may be degraded with over a few thousand threads in total. FCFC is found to be faster than most (if not all) other public pair-counting codes for modern cosmological pair-counting applications.

79 ASTRONOMY AND ASTROPHYSICS↗

Atomically dispersed Pt single sites and nanoengineered structural defects enable a high electrocatalytic activity and durability for hydrogen evolution reaction and overall urea electrolysis

The scarcity and high cost of PGM electrocatalysts are the key bottleneck in the mass-scale commercialization of many electrolysis technologies. Bifunctional single-atom electrocatalysts (SACs) are promising alternatives for PGM electrocatalysts in next-generation electrolysis technologies because of their superior intrinsic activity and perfect atom utilization. Regulating the coordination environment of platinum atomic sites identifies their electrocatalytic performance. Therefore, exploring more appropriate supports could facilitate the construction of active and durable electrocatalysts with ultralow noble metal content. Herein, we report on a reliable approach for producing a novel type of SACs composed of atomically dispersed Pt active sites stabilized on defective NiCo layered double hydroxide (Pt/D-NiCo LDH) nanosheets as an ultralow-Pt hybrid electrocatalyst for hydrogen evolution reaction (HER), urea oxidation reaction (UOR), and full urea-water electrolysis. The optimized Pt 1 /D NiCo LDH-24 SAC displays a remarkable HER and UOR performance where it yields a current density of 10 mA cm -2 at 37 mV and 1.25 V vs. RHE for HER and UOR, respectively. Finally, symmetrical urea electrolyzer constructed of Pt 1 /D-NiCo LDH-24 electrodes attains 10 mA cm -2 at a cell voltage of 1.32 V vs. RHE, demonstrating superior activity and durability over 60 h operation when compared to commercial Pt/C( + )||RuO 2 ( - ) system.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine learning for nonintrusive model order reduction of the parametric inviscid transonic flow past an airfoil

Fluid flow in the transonic regime finds relevance in aerospace engineering, particularly in the design of commercial air transportation vehicles. Computational fluid dynamics models of transonic flow for aerospace applications are computationally expensive to solve because of the high degrees of freedom as well as the coupled nature of the conservation laws. While these issues pose a bottleneck for the use of such models in aerospace design, computational costs can be significantly minimized by constructing special, structure-preserving surrogate models called reduced-order models. In this work, we propose a machine learning method to construct reduced-order models via deep neural networks and we demonstrate its ability to preserve accuracy with a significantly lower computational cost. In addition, our machine learning methodology is physics-informed and constrained through the utilization of an interpretable encoding by way of proper orthogonal decomposition. Application to the inviscid transonic flow past the RAE2822 airfoil under varying freestream Mach numbers and angles of attack, as well as airfoil shape parameters with a deforming mesh, shows that the proposed approach adapts to high-dimensional parameter variation well. Notably, the proposed framework precludes the knowledge of numerical operators utilized in the data generation phase, thereby demonstrating its potential utility in the fast exploration of design space for diverse engineering applications. Comparison against a projection-based nonintrusive model order reduction method demonstrates that the proposed approach produces comparable accuracy and yet is orders of magnitude computationally cheap to evaluate, despite being agnostic to the physics of the problem.

97 MATHEMATICS AND COMPUTING↗

Discrete-Event Model of WIPP Operations

The Waste Isolation Pilot Plant (WIPP) is the critical component of the Department of Energy's (DOE) Transuranic Radioactive Waste (TRU) disposition infrastructure. Quantifying the operations ofWIPP with a discrete-event model demonstrates the capability to assess efficiency, identity bottlenecks, and improve future operations. Such a model has been developed with the simulation software ExtendSim. This report outlines the structure of that model, summarizes the model's successful reproduction of the annual amount of emplaced waste, and demonstrates that WIPP is successfully receiving and emplacing waste at a rate consistent with the rate at which the waste arrives. The model serves as a first step toward future production enhancements at WIPP. Those enhancements will rely on close collaboration with the Carlsbad Field Office (CBFO), accurate interpretation and incorporation of the model results, and effective planning with other DOE Environmental Management (DOE-EM) entities.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Evolution of Metastable Phases During Mg Metal Corrosion: An In Situ Cryogenic X-ray Photoelectron Spectroscopy Study

Magnesium and its alloys are potential structural materials candidates for a wide variety of applications due to their high strength-to-weight ratio. However, ductility and poor corrosion resistance under ambient environmental conditions are the bottleneck for industrial deployment. Designing passivation layers and/or corrosion resistant alloys requires fundamental understanding of the corrosion process. The traditional ex-situ spectroscopic measurements of polycrystalline metal surface with ubiquitous surface impurities and grain boundaries only provided an indistinct view of the corrosion process. To clearly distinguish the mechanism and sequence of corrosion process, we employed in-situ cryo-based x-ray photoelectron spectroscopy (XPS) measurements on Mg single crystal surface exposed to aqueous salt solution. Clean Mg (0001) surfaces were exposed to pure D2O and NaCl aqueous solution (5 wt% NaCl+95 wt% D2O). The interfacial reactions were studied using multimodal analysis including XPS, x-ray diffraction (XRD) and scanning electron microscopy (SEM). In contrast to previous studies, our experiments demonstrated the formation of magnesium chloride hydroxide hydrate during aqueous salt corrosion processes. Evidence of metastable ClO* radicals were also found during the initial aqueous salt solution exposure.

Shutthanandan, Vaithiyalingam↗

Top Stack Optimization for Cu 2 BaSn(S, Se) 4 Photovoltaic Cell Leads to Improved Device Power Conversion Efficiency beyond 6%

Earth-abundant and air-stable Cu 2 BaSnS 4-x Se x (CBTSSe) and related thin-film absorbers are regarded as prospective options to meet the increasing demand for low-cost solar cell deployment. Devices based on vacuum-deposited CBTSSe absorbers have achieved record power conversion efficiency (PCE) of 5.2 % based on a conventional device structure using CdS buffer and i-ZnO/ITO window layers, with open-circuit voltage (V OC ) posing the major bottleneck for improving solar cell performance. The current study demonstrates a >20 % improvement in V OC (from 0.62 V to 0.75 V) and corresponding enhancement in PCE (from 5.1 % to 6.2 % without anti-reflection coating; to 6.5 % with MgF 2 anti-reflection coating) for solution-deposited CBTSSe solar cells. This performance improvement is realized by introducing an alternative successive ionic layer adsorption and reaction (SILAR)-deposited Zn 1-x Cd x S buffer combined with sputtered Zn 1-x Mg x O/Al-doped ZnO window/top contact layer, which offer lower electron affinities relative to the conventional CdS/i-ZnO/ITO stack and better matching with the low electron affinity of CBTSSe. A combined experimental (temperature- and light intensity-dependent V OC measurements) and device simulation (SCAPS-1D) evaluation points to the importance of addressing relative band offsets for both the buffer and window layers relative to the absorber in mitigating interfacial recombination and optimizing CBTSSe solar cell performance.

25 ENERGY STORAGE↗

A simulation framework for evaluating electronic order workflows in integrated health records

Electronic health record (EHR) systems are critical to modern healthcare delivery, yet the dynamic workflows that govern electronic order processing remain underexplored. Inefficiencies in these digital pathways can cause delays in care, repetitive workloads, and even patient harm. This study presents a discrete-event simulation framework used to reconstruct and evaluate EHR-based order workflows in a large integrated healthcare system. Using real-world data extracted from the Veterans Health Administration’s Corporate Data Warehouse, the authors mapped order events to standardized state transitions and modeled their progression across different facilities of varying complexity levels. After being calibrated with empirical distributions of transition times and validated against observed time-in-system metrics, the simulation demonstrates close alignment with historical performance. Scenario analyses reveal that resource capacity constraints significantly amplify the impact of electronic order surges, which are reflected in the disproportionate growth in backlogs and processing delays. Adjustments in transition probabilities further increased recirculation and extended workflow paths. Network-based analysis identified Reserved, InProgress, and Completed as structurally critical states that function as hubs within the process network but the transitions in-between also act as major bottlenecks. These results showcased the effectiveness of simulation-based approaches in monitoring EHR order processing performance and evaluating consequences of workflow changes on healthcare network resources planning. The proposed simulation framework provides a scalable data-driven tool to support operational decision-making and improve the efficiency of electronic order management in complex healthcare environments.

Engineering↗

A graph neural network-state predictive information bottleneck (GNN-SPIB) approach for learning molecular thermodynamics and kinetics

Molecular dynamics simulations offer detailed insights into atomic motions but face timescale limitations. Enhanced sampling methods have addressed these challenges but even with machine learning, they often rely on pre-selected expert-based features. Here, in this work, we present a Graph Neural Network-State Predictive Information Bottleneck (GNN-SPIB) framework, which combines graph neural networks and the state predictive information bottleneck to automatically learn low-dimensional representations directly from atomic coordinates. Tested on three benchmark systems, our approach predicts essential structural, thermodynamic and kinetic information for slow processes, demonstrating robustness across diverse systems. The method shows promise for complex systems, enabling effective enhanced sampling without requiring pre-defined reaction coordinates or input features.

Zou, Ziyue↗

Portability for GPU-accelerated molecular docking applications for cloud and HPC: can portable compiler directives provide performance across all platforms?

High-throughput structure-based screening of drug-like molecules has become a common tool in biomedical research. Recently, acceleration with graphics processing units (GPUs) has provided a large performance boost for molecular docking programs. Both cloud and high-performance computing (HPC) resources have been used for large screens with molecular docking programs; while NVIDIA GPUs have dominated cloud and HPC resources, new vendors such as AMD and Intel are now entering the field, creating the problem of software portability across different GPUs. Ideally, software productivity could be maximized with portable programming models that are able to maintain high performance across architectures. While in many cases compiler directives have been used as an easy way to offload parallel regions of a CPU-based program to a GPU accelerator, they may also be an attractive programming model for providing portability across different GPU vendors, in which case the porting process may proceed in the reverse direction: from low-level, architecture-specific code to higher-level directive-based abstractions. MiniMDock is a new mini-application (miniapp) designed to capture the essential computational kernels found in molecular docking calculations, such as are used in phar-maceutical drug discovery efforts, in order to test different solutions for porting across GPU architectures. Here we extend MiniMDock to GPU offloading with OpenMP directives, and compare to performance of kernels using CUDA and HIP on NVIDIA and AMD GPUs, respectively, as well as across different compilers, exploring performance bottlenecks. We document this reverse-porting process, from highly optimized device code to a higher-level version using directives, compare code structure, and describe barriers that were overcome in this effort.

Thavappiragasam, Mathialakan↗

End-to-end orientation estimation from 2D cryo-EM1images

Cryo-electron microscopy (cryo-EM) is a Nobel Prize-winning technique for deter-mining high-resolution 3D structures of biological macromolecules. A 3D structure is reconstructed from hundreds of thousands of noisy 2D projection images. However, existing 3D reconstruction methods are still time-consuming, and one of the major computational bottlenecks is to recover the unknown orientation of the particle in16each 2D image. The dominant methods typically exploit expensive global search on each image to estimate the missing orientations. Here, a novel end-to-end supervised learning method is introduced to directly recover the missing orientations from 2D cryo-EM images. A neural network is used to approximate the mapping from images to orientations. Furthermore, a robust loss function is proposed for optimizing the parameters of the network, which can handle both asymmetric and symmetric 3D structures. Experiments on synthetic datasets with various symmetry types confirm that the neural network is capable of recovering orientations from 2D cryo-EM images, and the results on one real cryo-EM dataset further demonstrate its potential in more challenging imaging conditions.

3D reconstruction↗

Structured illumination with thermal imaging (SI-TI): A dynamically reconfigurable metrology for parallelized thermal transport characterization

The recent push for the “materials by design” paradigm requires synergistic integration of scalable computation, synthesis, and characterization. Among these, techniques for efficient measurement of thermal transport can be a bottleneck limiting the experimental database size, especially for diverse materials with a range of roughness, porosity, and anisotropy. Traditional contact thermal measurements have challenges with throughput and the lack of spatially resolvable property mapping, while non-contact pump-probe laser methods generally need mirror smooth sample surfaces and also require serial raster scanning to achieve property mapping. Here, we present structured illumination with thermal imaging (SI-TI), a new thermal characterization tool based on parallelized all-optical heating and thermometry. Experiments on representative dense and porous bulk materials as well as a 3D printed thermoelectric thick film (~50 μm) demonstrate that SI-TI (1) enables paralleled measurement of multiple regions and samples without raster scanning; (2) can dynamically adjust the heating pattern purely in software, to optimize the measurement sensitivity in different directions for anisotropic materials; and (3) can tolerate rough (~3 μm) and scratched sample surfaces. Here, this work highlights a new avenue in adaptivity and throughput for thermal characterization of diverse materials.

42 ENGINEERING↗

Surface lattice engineering for fine-tuned spatial configuration of nanocrystals

Hybrid nanocrystals combining different properties together are important multifunctional materials that underpin further development in catalysis, energy storage, et al., and they are often constructed using heterogeneous seeded growth. Their spatial configuration (shape, composition, and dimension) is primarily determined by the heterogeneous deposition process which depends on the lattice mismatch between deposited material and seed. Precise control of nanocrystals spatial configuration is crucial to applications, but suffers from the limited tunability of lattice mismatch. Here, we demonstrate that surface lattice engineering can be used to break this bottleneck. Surface lattices of various Au nanocrystal seeds are fine-tuned using this strategy regardless of their shape, size, and crystalline structure, creating adjustable lattice mismatch for subsequent growth of other metals; hence, diverse hybrid nanocrystals with fine-tuned spatial configuration can be synthesized. This study may pave a general approach for rationally designing and constructing target nanocrystals including metal, semiconductor, and oxide.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Facilitating Machine Learning Collaborations Between Labs, Universities, And Industry

It is clear from numerous recent community reports, papers, and proposals that machine learning is of tremendous interest for particle accelerator applications. The quickly evolving landscape continues to grow in both the breadth and depth of applications including physics modeling, anomaly detection, controls, diagnostics, and analysis. Consequently, laboratories, universities, and companies across the globe have established dedicated machine learning (ML) and data science efforts aiming to make use of these new state-of-the-art tools. The current funding environment in the U.S. is structured in a way that supports specific application spaces rather than larger collaboration on community software. Here, we discuss the existing collaboration bottlenecks and how a shift in the funding environment, and how we develop collaborative tools, can help fuel the next wave of ML advancements for particle accelerators.

Edelen, J.P.↗