Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hierarchical optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Dual Channel Dual Staging: Hierarchical and Portable Staging for GPU-Based In-Situ Workflow

In-situ workflows have emerged as an attractive approach for addressing data movement challenges at very large scales. Since GPU-based architectures dominate the HPC landscapes, porting these in-situ workflows, and, specifically, the inter-application data exchange, to GPU-based systems can be challenging. Technologies such as GPUDirect RDMA (GDR), which is typically used for I/O in GPU applications as an optimization that circumvents the CPU overhead, can be leveraged to support bulk data exchanges between GPU applications. However, current GDR design often lacks performance portability across HPC clusters built with different hardware configurations. Furthermore, the local CPU may also be effectively used as an auxiliary communication mechanism to offload data exchanges. In this paper, we present a dual channel dual staging approach for efficient, scalable, and performance-portable inter-application data exchange for in-situ workflows. This approach exploits the data access pattern within in-situ workflows along with the inherent execution asynchrony to accelerate data exchanges and, at the same time, improve performance portability. Specifically, the dual channel dual staging method leverages both the local CPU and the remote data staging server to build a hierarchical joint staging area and uses this staging area to transform blocking inter-application bulk data exchanges into best-effort local data movements between GPU and CPU. The dual channel dual staging is implemented as a portability extension of the Dataspaces-GPU staging framework. We present an experimental evaluation of its performance, portability, and scalability using this implementation on three leadership GPU clusters. The evaluation results demonstrate that the dual channel dual staging method saves up to 75% in data-exchange time compared to host-based, GDR, and alternate portable designs, while maintaining scalability (up to 512 GPUs) and performance portability across the three platforms.

Zhang, Bo [University of Utah]↗

A Method for Producing Hierarchical and Statistically Calibrated Predictions of Nuclear Material Properties from Existing Models

Computer vision-based analysis of micrographs of nuclear materials is an emerging technique for property prediction, synthetic route identification, and other material analysis tasks. These analysis tasks play a pivotal role in many material characterization applications such as signature development for treaty verification, process optimization, etc. The backbone in many of the recent computer vision-based techniques is a deep learning model, which takes a fixed-size set of pixels and provides a class prediction for that set of pixels. For example, previous work developed a deep convolutional neural network (CNN) to predict the synthetic route from a 256 px x 256 px patch taken from a larger image of uranium ore concentrates. In this work, we present several methods for first calibrating these models in a manner that they can provide accurate probabilities of their predictions’ veracity, and several methods of combining these probabilities. Overall, the combination of these two steps into a pipeline allows for full-image and even full-sample (where a sample has many images) predictions with associated confidence values. Finally, we show that one can also use the patch predictions and confidence to produce a visualization to map predicted constituents through the image. Results and examples for predicting and mapping uranium ore concentrates’ synthetic process from imagery will be presented.

artificial intelligence↗

H3 Geospatial Mapping Resolution Recommendations

This report provides a brief overview of the Hexagonal Hierarchical Geospatial Indexing System (H3) and its benefits and use cases; a comparison of population estimates of various H3 resolutions with administrative boundaries (county, zip code, etc.); and H3 resolution recommendations for electric outage data reporting for US states considering optimal balance between accuracy, computational efficiency, and privacy preservation.

99 GENERAL AND MISCELLANEOUS↗

Architectural scaling tradeoffs in modular 3D bosonic quantum processors

We propose a modular three-dimensional bosonic quantum processor built from repeatable coupled-cavity modules linked by configurable interconnect networks. Using hardware-motivated graph-theoretic measures, we compare nearest-neighbor, hub-based, and hybrid architectures in terms of interconnect count, communication distance, resource concentration, and implementation complexity. Rather than identifying a universally optimal topology, our analysis shows how these architectures redistribute the costs of scaling, including wiring and port requirements, nonlocal communication distance, exposure to shared resources, routing bottlenecks, and scheduling overhead. Case studies of a \(3\times3\) processor and a larger hierarchical architecture further distinguish finite-size performance from asymptotic scaling. The resulting framework provides a systematic basis for evaluating modular three-dimensional bosonic processors and for identifying the device-level parameters required for quantitative hardware design.

Zhu, Shaojiang [Fermilab] (ORCID:0000000293180092)↗

PANDORA: A Parallel Dendrogram Construction Algorithm for Single Linkage Clustering on GPU

This paper introduces Pandora, a parallel algorithm for computing dendrograms, the hierarchical cluster trees for single linkage clustering (SLC). Current parallel approaches construct dendrograms by partitioning a minimum spanning tree and removing edges. However, they struggle with skewed, hard-to-parallelize real-world dendrograms. Consequently, computing dendrograms is the sequential bottleneck in HDBSCAN*[21], a popular SLC variant. Pandora uses recursive tree contraction to address this limitation. Pandora contracts nodes to construct progressively smaller trees. It computes the smallest contracted dendrogram and expands it by inserting contracted edges. This recursive strategy is highly parallel, skew-independent, work-optimal, and well-suited for GPUs and multicores. We develop a performance portable implementation of Pandora in Kokkos[31] and evaluate its performance on multicore CPUs and multi-vendor GPUs (e.g., Nvidia, AMD) for dendrogram construction in HDBSCAN*. Multithreaded Pandora is 2.2x faster than the current best-multithreaded implementation. Our GPU version achieves 6-20x speedup on AMD GPUs and 10-37x on NVIDIA GPUs over multithreaded Pandora. Pandora removes HDBSCAN*’s sequential bottleneck, greatly boosting efficiency, particularly with GPUs.

Sao, Piyush↗

Constructing Highly Porous Low Iridium Anode Catalysts Via Dealloying for Proton Exchange Membrane Water Electrolyzers

Iridium (Ir) is the most active and durable anode catalyst for the oxygen evolution reaction (OER) for proton exchange membrane water electrolyzers (PEMWEs). However, their large-scale applications are hindered by high costs and scarcity of Ir. Lowering Ir loadings below 1.0 mgcm -2 causes significantly reduced PEMWE performance and durability. Therefore, developing efficient low Ir-based catalysts is critical to widely commercializing PEMWEs. Herein, an approach is presented for designing porous Ir metal aerogel (MA) catalysts via chemically dealloying IrCu alloys. In this study, the unique hierarchical pore structures and multiple channels of the Ir MA catalyst significantly increase electrochemical surface area (ECSA) and enhance OER activity compared to conventional Ir black catalysts, providing an effective solution to design low-Ir catalysts with improved Ir utilization and enhanced stability. An optimized membrane electrode assembly (MEA) with an Ir loading of 0.5 mg Ir cm -2 generated 2.0 A cm -2 at 1.79 V, higher than the Ir black at a loading of 2.0 mg Ir cm -2 (1.63 A cm -2 ). The low-Ir MEA demonstrated an acceptable decay rate of ≈40 µV h -1 during durability tests at 0.5 (>1200 h) and 2.0 A cm -2 (400 h), outperforming the commercial Ir-based MEA (175 µV h -1 at 2.0 mg Ir cm -2 ).

36 MATERIALS SCIENCE↗

Grid-Aware Charging and Operational Optimization for Mixed-Fleet Public Transit

The rapid growth of urban populations and the increasing need for sustainable transportation solutions have prompted a shift towards electric buses in public transit systems. However, the effective management of mixed fleets consisting of both electric and diesel buses poses significant operational challenges. One major challenge is coping with dynamic electricity pricing, where charging costs vary throughout the day. Transit agencies must optimize charging assignments in response to such dynamism while accounting for secondary considerations such as seating constraints. This paper presents a comprehensive mixed-integer linear programming (MILP) model to address these challenges by jointly optimizing charging schedules and trip assignments for mixed (electric and diesel bus) fleets while considering factors such as dynamic electricity pricing, vehicle capacity, and route constraints. We address the potential computational intractability of the MILP formulation, which can arise even with relatively small fleets, by employing a hierarchical approach tailored to the fleet composition. By using real-world data from the city of Chattanooga, Tennessee, USA, we show that our approach can result in significant savings in the operating costs of the mixed transit fleets.

Sen, Rishav↗

An experimentally informed design process for future inertial confinement fusion facilities

The achievement of ignition in the laboratory has renewed interest in defining the requirements for a future high-gain inertial confinement fusion (ICF) facility. Our best chance of predicting future ICF performance is with 3-D radiation hydrodynamic simulations that have been benchmarked against experimental data, but their high computational cost is prohibitive for use in practical design studies. We introduce a hierarchical approach where 3-D simulations are tuned to match experimental measurements and used to train 3-D degradation models in 1-D simulations allowing for accurate predictions over the entire OMEGA direct-drive database. A genetic algorithm was used in combination with the trained 1-D simulations to search for optimal direct-drive implosion designs at driver energies ranging from 20 kJ to 10 MJ. As the fidelity of 3-D codes improves, this approach will provide a viable experimentally informed tool for defining the next ICF facility.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

CSGL: chemical synthesis graph learning for molecule representation

Abstract Motivation Molecule representation learning (MRL) translates molecules into a real vector space, serving as input to downstream tasks in biology, chemistry, and computer science. This article introduces a chemical synthesis graph learning (CSGL) framework, which enhances MRL by considering both the atomic structures of molecules and their roles in chemical reactions through a hierarchical graph representation. Specifically, molecules are first modeled based on their molecular graphs, which capture atomic-level structural information. They are then further refined using a chemical synthesis graph, where nodes represent reactant and product molecule sets, and edges encode chemical transformations between reactants and products (e.g. changes in molecular structures). CSGL optimizes molecular embeddings of reactant and product nodes in a fashion that ensures the embeddings conform to a chemical balance constraint. Results Experimental results show that our method CSGL achieves strong performance on a variety of tasks, including product prediction, reaction classification, and molecular property prediction. Availability and implementation https://github.com/li-2023/CSGL.

Biochemistry & Molecular Biology↗

Harnessing Virtual Power Plants Reliably: Enabling tools for increased observability, controllability, operation, and aggregation of distributed energy resources

Harnessing virtual power plants enhances the integration of distributed energy resources into utility grids for a sustainable energy future. Virtual power plants (VPPs) aggregate DERs to enhance resource adequacy and reduce emissions. U.S. utilities are exploring various technologies to manage DERs effectively. FERC Order 2222 allows DERs to participate in both wholesale and retail markets. Enhancing observability and controllability of behind-the-meter (BTM) DERs is essential for reliable grid operations. A hierarchical control architecture can improve coordination among residential energy resources. Field tests showed nearly 20% energy savings and 30% peak power reduction during grid events. Effective management of DERs requires enhanced situational awareness to prevent grid congestion. Integrating DER management systems (DERMS) with existing planning tools can improve operational security. Near-real-time grid models can validate optimal resource set points against resource uncertainty. Traditional uninterruptible power supplies (UPS) can be upgraded to support grid services and become part of VPPs. Upgrading UPS systems can reduce costs by 75% and unlock significant battery capacity. New battery management systems and grid-aware controllers are essential for optimizing UPS performance. Continued research and development are necessary to address challenges in integrating DERs into utility grids. Encouraging customer participation in pilot programs is vital for the evolution of VPPs. Here, the shift towards price-responsive DERs and VPPs is expected to enhance energy distribution efficiency.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Tree tensor network hierarchical equations of motion based on time-dependent variational principle for efficient open quantum dynamics in structured thermal environments

In this work, we introduce an efficient method, TTN-HEOM, for exactly calculating the open quantum dynamics for driven quantum systems interacting with highly structured bosonic baths by combining the tree tensor network (TTN) decomposition scheme with the bexcitonic generalization of the numerically exact hierarchical equations of motion (HEOM). The method yields a series of quantum master equations for all core tensors in the TTN that efficiently and accurately capture the open quantum dynamics for non-Markovian environments to all orders in the system–bath interaction. These master equations are constructed based on the time-dependent Dirac–Frenkel variational principle, which isolates the optimal dynamics for the core tensors given the TTN ansatz. The dynamics converges to the HEOM when increasing the rank of the core tensors, a limit in which the TTN ansatz becomes exact. We introduce TENSO, tensor equations for non-Markovian structured open systems, as a general-purpose Python code to propagate the TTN-HEOM dynamics. We implement three general propagators for the coupled master equations: two fixed-rank methods that require a constant memory footprint during the dynamics and one adaptive-rank method with a variable memory footprint controlled by the target level of computational error. We exemplify the utility of these methods by simulating a two-level system coupled to a structured bath containing one Drude–Lorentz component and eight Brownian oscillators, which is beyond what can presently be computed using the standard HEOM. Our results show that the TTN-HEOM is capable of simulating both dephasing and relaxation dynamics of driven quantum systems interacting with structured baths, even those of chemical complexity, with an affordable computational cost.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Breaking the curse of dimensionality: Solving configurational integrals for crystalline solids by tensor networks

Accurately evaluating configurational integrals for dense solids remains a central and difficult challenge in the statistical mechanics of condensed systems. Here, we present a tensor network approach that reformulates the high-dimensional configurational integral for identical-particle crystals into a sequence of computationally efficient summations. We represent the integrand as a high-dimensional tensor and apply tensor-train (TT) decomposition together with a custom TT-cross interpolation. This approach circumvents the need to explicitly construct the full tensor. We introduce tailored rank-1 and rank-2 schemes optimized for sharply peaked Boltzmann probability densities, typical for identical-particle crystals. When applied to the calculation of internal energy and pressure-temperature curves for crystalline Cu and Ar at high (GPa) pressures, as well as the alpha-to-beta phase transition diagram of Sn, our method accurately reproduces molecular dynamics simulation results using tight-binding, machine learning, hierarchical interacting particle–neural network, and modified embedded atom method potentials,all within seconds of computation time.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

New Era Towards Autonomous Additive Manufacturing: A Review of Recent Trends and Future Perspectives

Abstract The Additive Manufacturing (AM) landscape has significantly transformed in alignment with Industry 4.0 principles, primarily driven by the integration of Artificial Intelligence (AI) and Digital Twin (DT). However, current Intelligent Additive Manufacturing (IAM) systems face limitations such as fragmented AI tool usage and suboptimal human-machine interaction (HMI). This paper reviews existing IAM solutions, emphasizing control, monitoring, process autonomy, and end-to-end integration, and identifies key limitations, such as the absence of a high-level controller for global decision-making. To address these gaps, we propose a transition from IAM to Autonomous Additive Manufacturing (AAM), featuring a hierarchical framework with four integrated layers: knowledge, generative solution, operational, and cognitive. In the cognitive layer, AI agents notably enable machines to independently observe, analyze, plan, and execute operations that traditionally require human intervention. These capabilities streamline production processes and expand the possibilities for innovation, particularly in sectors like in-space manufacturing (ISM). Additionally, this paper discusses the role of AI in self-optimization and lifelong learning, positing that the future of AM will be characterized by a symbiotic relationship between human expertise and advanced autonomy, fostering a more adaptive, resilient manufacturing ecosystem.

Fan, Haolin↗

Laboratory Evaluation of Federated, Hierarchical Controls for Distribution Power System Management: Preprint

The connection of more loads and distributed energy resources (DERs) to the distribution power system brings both challenges and opportunities to system operators. There are opportunities to aggregate flexible loads and DERs to provide transmission grid services, but the coordinated actions of DERs being managed by independent, third-party DER aggregators to support transmission system operations can present challenges. We developed a federated DER management architecture and control framework that aims to manage heterogeneous DERs to deliver reliable transmission grid services while respecting distribution system constraints. The controls include stochastic day-ahead optimization, model predictive control, and a simple real-time management scheme. We present simulation results obtained from a realistic laboratory test bed of federated controls managing DERs within a substation service area to make the substation net power follow the optimal net power determined by the day-ahead optimization based on cost and limiting reverse power flow.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Flexible Resource Scheduler for FAST-DERMS (FRS-FASTDERMS) v0.9

The Flexible Resource Scheduler is a hierarchical controller that manages the distributed energy resources in a distribution substation or distribution feeder to provide a firm commitment of power flow at the substation or feeder head to be scheduled in transmission-level markets as an aggregated demand resource. It is the reference controller for the FAST-DERMS Architecture, developed in tandem with the architecture under the DOE FAST-DERMS project. It is comprised of a day-ahead stochastic optimization, which schedules substation power flow and reserves, a intra-hour MPC, which generates dispatch base points for DER, and a real-time PID controller maintaining that dispatches DER to maintain the substation power around the base points. The repository also includes a representative aggregator controller, and all of the necessary components to run a simulation using PNNL's GridAPPS-D software with the controller.

MacDonald, Jason [Lawrence Berkeley National Labor↗

Uncovering multiscale structure-property correlations via active learning in scanning tunneling microscopy

Atomic arrangements and local sub-structures fundamentally influence emergent material functionalities. These structures are conventionally probed using spatially resolved studies and the property correlations are deciphered by a researcher based on sequential explorations, thereby limiting the efficiency and scope. Here we demonstrate a multi-scale Bayesian deep-learning based framework that automatically correlates material structure with its electronic properties using scanning tunneling microscopy (STM) measurements in real-time. Its predictions are used to autonomously direct exploration toward regions of the sample that optimize a given material property. This method is deployed on a low-temperature ultra-high vacuum STM to understand the structure-property relationship in a europium-based semimetal, EuZn 2 As 2 , a promising candidate relevant to magnetism-driven topological phenomena. The framework employs a sparse-sampling approach to efficiently construct the scalar-property space using minimal measurements, about 1–10% of the data required in standard hyperspectral methods. Moreover, we formulate the problem hierarchically across length scales, implementing autonomous workflow to locate mesoscopic and atomic structures that correspond to a target material property. This framework offers the choice to design scalar-property from the spectroscopic data to steer sample exploration. Our findings reveal correlations of the electronic properties unique to surface terminations, local defect density, and point defects.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Hydrogen Bond Benchmark: Focal‐Point Analysis and Assessment of DFT Functionals

We performed a hierarchical, convergent ab initio benchmark study and systematically analyzed the performance of density functional approximations for describing hydrogen bonds in small neutral, cationic, and anionic complexes, as well as in larger systems involving amide, urea, deltamide, and squaramide moieties. Focal point analyses (FPA), extrapolating to the ab initio limit, were carried out using correlated wave function methods up to CCSDT(Q) for the small complexes and CCSD(T) for the larger systems, together with correlation-consistent Gaussian basis sets up to the complete basis set limit. Optimized geometries and vibrational frequencies were obtained at the CCSD(T) level. The resulting FPA hydrogen-bond energies converge within a few tenths of a kcal mol −1 . These reference data were used to evaluate 60 density functionals (including 12 dispersion-corrected), spanning the local-density approximation (LDA), generalized gradient approximations (GGAs), meta-GGAs, hybrids, meta-hybrids, double-hybrids, and range-separated hybrids. Overall, the meta-hybrid M06-2X provides the best performance for both hydrogen bond energies and geometries, while the dispersion-corrected GGAs BLYP-D3(BJ) and BLYP-D4 also yield accurate hydrogen-bond data and can serve as cost-effective options for studying large and complex systems.

coupled cluster theory↗

Hierarchical Speed Planner for Automated Vehicles: A Framework for Lagrangian Variable Speed Limit in Mixed-Autonomy Traffic

Here, this article presents a novel hierarchical speed planning framework for variable speed limits in mixed-autonomy traffic environments, leveraging server-side macroscopic control and vehicle-side microscopic execution. The framework integrates real-time traffic state estimation (TSE) and reinforcement learning (RL)-based control to mitigate congestion and improve traffic flow. A TSE enhancement module combines macroscopic data from sources like INRIX with high-resolution observations from connected autonomous vehicles (CAVs), enabling predictive modeling to address latency and noise. The target speed design module employs kernel smoothing and a buffer zone strategy to optimize traffic density and flow around bottlenecks. The proposed system was validated in the largest open-road test to date with 100 CAVs, demonstrating an overall 8% traffic density decrease, with a specific decrease of 7% upstream, 10% downstream, and a 52% decrease during the congestion formation phase at bottlenecks.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗