Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “partitioned algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Probabilistic partition of unity networks for high–dimensional regression problems

We explore the probabilistic partition of unity network (PPOU-Net) model in the context of high-dimensional regression problems and propose a general framework focusing on adaptive dimensionality reduction. With the proposed framework, the target function is approximated by a mixture of experts model on a low-dimensional manifold, where each cluster is associated with a fixed-degree polynomial. We present a training strategy that leverages the expectation maximization (EM) algorithm. During the training, we alternate between (i) applying gradient descent to update the DNN coefficients; and (ii) using closed-form formulae derived from the EM algorithm to update the mixture of experts model parameters. Under the probabilistic formulation, step (ii) admits the form of embarrassingly paralleliazable weighted least-squares solves. The PPOU-Nets consistently outperform the baseline fully-connected neural networks of comparable sizes in numerical experiments of various data dimensions. Here, we also explore the proposed model in applications of quantum computing, where the PPOU-Nets act as surrogate models for cost landscapes associated with variational quantum circuits.

97 MATHEMATICS AND COMPUTING↗

Real-time streamflow forecasting: AI vs. Hydrologic insights

In this paper, we propose a set of simple benchmarks for the evaluation of data-based models for real-time streamflow forecasting, such as those developed with sophisticated Artificial Intelligence (AI) algorithms. The benchmarks are also data-based and provide context to judge incremental improvements in the performance metrics from the more complicated approaches. The benchmarks include temporal and spatial persistence, persistence corrected for baseflow and streamflow, as well as river distance weighted runoff obtained from space-time distributed rainfall. In the development of the benchmarks, we use basic hydrologic insights such as flow aggregation by the river network, scale-dependence in basin response, streamflow partitioning into quick flow and baseflow, water travel time, and rainfall averaging by the basin width function. The study uses 140 streamflow gauges in Iowa that cover a range of basin scales between 7 and 37,000 km 2 . The data cover 17 years. This work demonstrates that the proposed benchmarks can provide good performance according to several commonly used metrics. For example, streamflow forecasting at half of the test locations across years achieves a Kling-Gupta Efficiency (KGE) score of 0.6 or higher at one-day ahead lead time, and 20% of cases reach the KGE of 0.8 or higher. The proposed benchmarks are easy to implement and should prove useful for developers of data-based as well as physics-based hydrologic models and real-time data assimilation techniques.

54 ENVIRONMENTAL SCIENCES↗

Enhancing scalability of a matrix-free eigensolver for studying many-body localization

We propose several techniques to enhance the parallel scalability of a matrix-free eigensolver designed for studying many-body localization (MBL) of quantum spin chain models with nearest-neighbor interactions and on-site disorder. This type of problem is computationally challenging because the dimension of the associated Hamiltonian matrix grows exponentially with respect to the number of spins L, and we need to average over different realizations of the random disorder to obtain relevant statistical behavior. For each disorder realization, we need to compute eigenvalues from different regions of the spectrum and their corresponding eigenvectors. In previous work, the interior eigenstates for a single eigenvalue problem are computed via the shift-and-invert Lanczos algorithm. Due to the extremely high memory footprint of the LU factorizations, this technique is not well suited for large L’s. For example, we need thousands of compute nodes on modern high performance computing infrastructures to go beyond L = 24. The matrix-free approach does not suffer from this memory bottleneck, however, its scalability is limited by a computation and communication load imbalance. To reduce this imbalance and to significantly enhance the scalability of the matrix-free eigensolver, we reorder the matrix and leverage the consistent space runtime, CSPACER. We also show its efficiency in managing irregular communication patterns at scale compared to optimized MPI non-blocking two-sided and one-sided RMA implementation variants. This effort enables us to study MBL for spin chains with a larger number of spins. The efficiency and effectiveness of the proposed algorithm is demonstrated by computing eigenstates on a massively parallel many-core high performance computer.

METIS↗

Intelligent Partitioning based Fully Parallel AC Security-Constrained Optimal Power Flow

Today’s power grid is becoming more diverse and integrated with high-level distributed energy resources and smart control technologies that is creating a new set of grid management challenges in terms of large-scale, nonlinear, and non-convex problem modeling, complex and time-consuming computation, as well as difficult uncertainty handling. This project focused on solving a challenging multi-period security-constrained generation scheduling problem, which is of great importance for maximizing the social welfare of real-time dispatch, day-ahead market, as well as weekly planning of power systems. Our developed software explored parallel optimization algorithms for complex and realistic power system models, and develop fast, efficient, and robust grid optimization solutions on the high-performance computing platform that will enable increased grid economics, flexibility, resilience, as well as energy security in the United States.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Ab Initio Direct Dynamics

The reactivity and dynamics of molecular systems can be explored computationally by classical trajectory calculations. The traditional approach involves fitting a functional form of a potential energy surface (PES) to the energies from a large number of electronic structure calculations and then integrating numerous trajectories on this fitted PES to model the molecular dynamics. The ever-decreasing cost of computing and continuing advances in computational chemistry software have made it possible to use electronic structure calculations directly in molecular dynamics simulations without first having to construct a fitted PES. In this “on-the-fly” approach, every time the energy and its derivatives are needed for the integration of the equations of motion, they are obtained directly from quantum chemical calculations. This approach started to become practical in the mid-1990s as a result of increased availability of inexpensive computer resources and improved computational chemistry software. The application of direct dynamics calculations has grown rapidly over the last 25 years and would require a lengthy review article. The present Account is limited to some of our contributions to methods development and various applications. To improve the efficiency of direct dynamics calculations, we developed a Hessian-based predictor-corrector algorithm for integrating classical trajectories. Hessian updating made this even more efficient. Furthermore, this approach was also used to improve algorithms for following the steepest descent reaction paths. For larger molecular systems, we developed an extended Lagrangian approach in which the electronic structure is propagated along with the molecular structure. Strong field chemistry is a rapidly growing area, and to improve the accuracy of molecular dynamics in intense laser fields, we included the time-varying electric field in a novel predictor-corrector trajectory integration algorithm. Since intense laser fields can excite and ionize molecules, we extended our studies to include electron dynamics. Specifically, we developed code for time-dependent configuration interaction electron dynamics to simulate strong field ionization by intense laser pulses. Our initial application of ab initio direct dynamics in 1994 was to CH 2 O → H 2 + CO; the calculated vibrational distributions in the products were in very good agreement with experiment. In the intervening years, we have used direct dynamics to explore energy partitioning in various dissociation reactions, unimolecular dissociations yielding three fragments, reactions with branching after the transition state, nonstatistical dynamics of chemically activated molecules, dynamics of molecular fragmentation by intense infrared laser pulses, selective activation of specific dissociation channels by aligned intense infrared laser fields, angular dependence of strong field ionization, and simulation of sequential double ionization.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

OPER: Optimality-Guided Embedding Table Parallelization for Large-scale Recommendation Model

With the sharp increasing volume of user data, Deep Learning Recommendation Model (DLRM) becomes an indispensable infrastructure in large technology companies. However, large-scale DLRM on the multi-GPU platform is still inefficient due to unbalanced workload partitioning and intensive inter-GPU communication. To this end, we propose OPER, an OPtimality guided Embedding table placement for large-scale Recommendation model training and inference. OPER explores the potential of mitigating remote memory access latency in DLRM through fine-grained embedding table placement. Specifically, OPER proposes a theoretical modeling that builds up the relationship between EMT placement and the embedding communication latency in both training and inference. OPER proves the NP hardness of finding the optimal embedding table placement and proposes a heuristic algorithm that yields near optimal placement. OPER implements a SHMEM-based embedding table training system and a unified embedding index mapping to support fine-grained embedding table sharding and placement. Comprehensive experiments reveal that OPER achieves on average 3.4× and 5.1× speedup on training and inference respectively over state-of-the-art DLRM frameworks.

Wang, Zheng↗

A Modular and Transferable Reinforcement Learning Framework for the Fleet Rebalancing Problem

Mobility on demand (MoD) systems show great promise in realizing flexible and efficient urban transportation. However, significant technical challenges arise from operational decision making associated with MoD vehicle dispatch and fleet rebalancing. For this reason, operators tend to employ simplified algorithms that have been demonstrated to work well in a particular setting. To help bridge the gap between novel and existing methods, we propose a modular framework for fleet rebalancing based on model-free reinforcement learning (RL) that can leverage an existing dispatch method to minimize system cost. In particular, by treating dispatch as part of the environment dynamics, a centralized agent can learn to intermittently direct the dispatcher to reposition free vehicles and mitigate against fleet imbalance. We formulate RL state and action spaces as distributions over a grid partitioning of the operating area, making the framework scalable and avoiding the complexities associated with multiagent RL. Numerical experiments, using real-world trip and network data, demonstrate that RL reduces waiting time by 28% to 38% for the same-day evaluation, 17% to 44% for cross-day evaluation, and 22% to 25% for cross-season evaluation compared with no rebalancing scenarios. This approach has several distinct advantages over baseline methods including: improved system cost; high degree of adaptability to the selected dispatch method; and the ability to perform scale-invariant transfer learning between problem instances with similar vehicle and request distributions.

33 ADVANCED PROPULSION SYSTEMS↗

Dimensionally reduced machine learning model for predicting single component octanol–water partition coefficients

Abstract MF-LOGP, a new method for determining a single component octanol–water partition coefficients ( $$LogP$$ LogP ) is presented which uses molecular formula as the only input. Octanol–water partition coefficients are useful in many applications, ranging from environmental fate and drug delivery. Currently, partition coefficients are either experimentally measured or predicted as a function of structural fragments, topological descriptors, or thermodynamic properties known or calculated from precise molecular structures. The MF-LOGP method presented here differs from classical methods as it does not require any structural information and uses molecular formula as the sole model input. MF-LOGP is therefore useful for situations in which the structure is unknown or where the use of a low dimensional, easily automatable, and computationally inexpensive calculations is required. MF-LOGP is a random forest algorithm that is trained and tested on 15,377 data points, using 10 features derived from the molecular formula to make $$LogP$$ LogP predictions. Using an independent validation set of 2713 data points, MF-LOGP was found to have an average $$RMSE$$ RMSE = 0.77 ± 0.007, $$MAE$$ MAE = 0.52 ± 0.003, and $${R}^{2}$$ R 2 = 0.83 ± 0.003. This performance fell within the spectrum of performances reported in the published literature for conventional higher dimensional models ( $$RMSE$$ RMSE = 0.42–1.54, $$MAE$$ MAE = 0.09–1.07, and $${R}^{2}$$ R 2 = 0.32–0.95). Compared with existing models, MF-LOGP requires a maximum of ten features and no structural information, thereby providing a practical and yet predictive tool. The development of MF-LOGP provides the groundwork for development of more physical prediction models leveraging big data analytical methods or complex multicomponent mixtures. Graphical Abstract

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Performance Analysis of Speculative Parallel Adaptive Local Timestepping for Conservation Laws

Stable simulation of conservation laws, such as those used to model fluid dynamics and plasma physics applications, requires the satisfaction of the so-called Courant-Friedrichs-Lewy condition. By allowing regions of the mesh to advance with different timesteps that locally satisfy this stability constraint, significant work reduction can be attained when compared to a time integration scheme using a single timestep size. However, parallelizing this algorithm presents considerable difficulty. Since the stability condition depends on the state of the system, dependencies become dynamic and potentially non-local. In this article, we present an adaptive local timestepping algorithm using an optimistic (Timewarp-based) parallel discrete event simulation. We introduce waiting heuristics to limit misspeculation and a semi-static load balancing scheme to eliminate load imbalance as parts of the mesh require finer or coarser timesteps. Last, we outline an interface for separating the physics of the specific conservation law from the temporal integration allowing for productive adoption of our proposed algorithm. We present a misspeculation study for three conservation laws, demonstrating both the productivity of the local timestepping API, for which 74% of the lines of code are reused across different conservation laws, and the robustness of the waiting heuristics—at most 1.5% of element updates are rolled back. Our performance studies demonstrate up to a 2.8× speedup versus a baseline unoptimized local timestepping approach, a 4x improvement in per-node throughput compared to an MPI parallelization of synchronous timestepping, and scalability up to 3,072 cores on NERSC’s Cori Haswell partition.

97 MATHEMATICS AND COMPUTING↗

Intelligent Experiments through Real-Time AI: Fast Data Processing and Autonomous Detector Control for High-Energy Nuclear Experiments

The aim of this project is to develop software and hardware for fast real-time data processing and autonomous detector control and calibration for the sPHENIX and the future EIC experiments. Below summarizes Georgia Tech team efforts in the past year: 1. We developed a real-time clustering algorithm and FPGA-based pipeline architecture for processing fired pixel data from ALPIDE sensors in sPHENIX experiments. Our Columnar Clustering Co-Design introduces a hardware-aware, stream-friendly approach that segments pixel data by column pairs using a Column Pair Clustering (CPC) strategy, followed by Cluster Stitching to merge adjacent subclusters. Implemented in Vitis HLS, the pipeline comprises five stages—read-in, subclustering, stitching, analysis, and write-out—connected by tagged HLS streams with custom end-of-event signaling for robust synchronization. We designed a pipelined dataflow model optimized for throughput, low latency, and minimal buffering, enabling scalable clustering across events of arbitrary size. Our system maintains spatial precision via center-of-mass and shape key extraction and efficiently handles edge cases such as fragmented or nested clusters. Compared against DBSCAN in both software and hardware, our approach demonstrates competitive performance under FPGA constraints. 2. We also conducted a comprehensive algorithm-to-hardware co-design of connected component analysis tailored for sPHENIX experiments, focusing on real-time, low-latency processing using FPGAs and High-Level Synthesis (HLS). Starting from a Python-based particle tracking pipeline, the team translated the core logic—graph traversal via DFS and Union-Find—into an HLS-compatible C++ model, replacing dynamic memory and recursion with static arrays and pipelined control flow. The final design includes a fully streamed and dataflow-compatible Union-Find kernel optimized across five iterations, incorporating loop pipelining, array partitioning, AXI/FIFO interface tuning, and function flattening. Experimental results show up to 14.8× speedup over the CPU baseline, reducing per-graph latency to 1.58 μs and demonstrating strong resource efficiency with only ~7k LUTs and zero BRAM usage. The design maintains functional correctness against the Python reference using a Python-based C-simulation framework and Mean Squared Error metrics. This work validates the potential of HLS-driven FPGA designs for edge-level HEP data acquisition, laying a scalable foundation for future integration with real-time detector pipelines and multi-graph processing systems.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Decentralized modular hybrid supervisory control for the formation of unmanned helicopters

Abstract Formation control of Unmanned Aerial Vehicles (UAVs) requires them to tightly cooperate to reach and keep the formation, while avoiding collision. This paper proposes a novel decentralized hybrid supervisory control approach for the formation control of multiple UAVs. This is achieved by developing a symbolic motion planning technique to polarly partition the motion space resulting in a finite state discrete event model for the motion dynamics of each UAV. Then, a modular discrete supervisor is designed for different components of the formation mission including reaching the formation, keeping the formation, and collision avoidance. Further, for the collision avoidance mechanism, a novel top‐down decomposition‐based approach is developed to design local supervisors decentralizedly. It is formally proved that with the proposed top‐down decomposition‐based approach, the (locally) supervised UAVs, as a whole, can cooperatively satisfy the desired (global) collision avoidance specification. The proposed decentralized supervisory control algorithm is also verified through a hardware‐in‐the‐loop simulator for the formation control of unmanned helicopters.

Karimoddini, Ali↗

The impact of aerosol mixing state on immersion freezing: insights from classical nucleation theory and particle-resolved simulations

Immersion freezing, initiated by ice-nucleating particles (INPs) in supercooled aqueous droplets, plays an important role in the formation of ice crystals within clouds. The efficiency of immersion freezing depends strongly on INP composition and, crucially, on the mixing state – how chemical species are distributed across the particle population. Here, we quantify the impact of aerosol mixing state on immersion freezing using a combined theoretical and particle-resolved modeling approach. We derive analytical expressions for the frozen fraction of internally and externally mixed INP populations based on classical nucleation theory, showing that the frozen fraction is sensitive to whether ice-active species are present in all particles or only in a subset of the population. We introduce a multi-species immersion freezing scheme into the particle-resolved model PartMC, using the water activity-based immersion freezing model (ABIFM) to compute freezing probabilities for mixed-composition particles. To improve computational efficiency, we implement a Binned Tau-Leaping algorithm and demonstrate an order-of-magnitude speedup with minimal accuracy loss. Simulations reproduce the analytical trends in limiting cases and extend the analysis to more general aerosol populations, where mixing state continues to exert a substantial control on frozen fraction. Sensitivity analyses across particle size, species type, and cooling condition reveal that the mixing state effect is most pronounced when small amounts of highly efficient INPs are mixed with less efficient materials. These findings underscore the need to represent aerosol mixing state explicitly in models of heterogeneous ice nucleation to reduce uncertainty in cloud-phase partitioning.

54 ENVIRONMENTAL SCIENCES↗

A coupled discontinuous Galerkin-Finite Volume framework for solving gas dynamics over embedded geometries

Herein, we present a computational framework for solving the equations of inviscid gas dynamics using structured grids with embedded geometries. The novelty of the proposed approach is the use of high-order discontinuous Galerkin (dG) schemes and a shock-capturing Finite Volume (FV) scheme coupled via an hp adaptive mesh refinement (hp-AMR) strategy that offers high-order accurate resolution of the embedded geometries. The hp-AMR strategy is based on a multi-level block-structured domain partition in which each level is represented by block-structured Cartesian grids and the embedded geometry is represented implicitly by a level set function. The intersection of the embedded geometry with the grids produces the implicitly-defined mesh that consists of a collection of regular rectangular cells plus a relatively small number of irregular curved elements in the vicinity of the embedded boundaries. High-order quadrature rules for implicitly-defined domains enable high-order accuracy resolution of the curved elements with a cell-merging strategy to address the small-cell problem. The hp-AMR algorithm treats the system with a second-order finite volume scheme at the finest level to dynamically track the evolution of solution discontinuities while using dG schemes at coarser levels to provide high-order accuracy in smooth regions of the flow. On the dG levels, the methodology supports different orders of basis functions on different levels. The space-discretized governing equations are then advanced explicitly in time using high-order Runge-Kutta algorithms. Numerical tests are presented for two-dimensional and three-dimensional problems involving an ideal gas. The results are compared with both analytical solutions and experimental observations and demonstrate that the framework provides high-order accuracy for smooth flows and accurately captures solution discontinuities.

97 MATHEMATICS AND COMPUTING↗

PUMIPic: A mesh-based approach to unstructured mesh Particle-In-Cell on GPUs

Unstructured mesh particle-in-cell, PIC, simulations executing on the current and next generation of massively parallel systems require new methods for both the mesh and particles to achieve performance and scalability on GPUs. The traditional approach to implementing PIC simulations defines data structures and algorithms in terms of particles with a full copy of the unstructured mesh on every process. To effectively scale the unstructured mesh and particles, mesh-based PIC uses the unstructured mesh as the predominant data structure with the particles stored in terms of the mesh entities. Here, this paper details the PUMIPic library, a framework for developing efficient and performance-portable mesh-based PIC simulations on GPU systems. A pseudo physics simulation based on a five-dimensional gyro-kinetic code for modeling plasma physics is used to examine the performance of PUMIPic. Scaling studies of the unstructured mesh partition and number of particles are performed up to 4096 nodes of the Summit system at Oak Ridge National Laboratory. The studies show that mesh-based PIC can utilize a partitioned mesh and maintain scaling up to system limitations.

97 MATHEMATICS AND COMPUTING↗

Application of the variational autoencoder to detect the critical points of the anisotropic Ising model

We generalize the previous study on the application of variational autoencoders to the two-dimensional Ising model to a system with anisotropy. Due to the self-duality property of the system, the critical points can be located exactly for the entire range of anisotropic coupling. This presents an excellent test bed for the validity of using a variational autoencoder to characterize an anisotropic classical model. Furthermore, we reproduce the phase diagram for a wide range of anisotropic couplings and temperatures via a variational autoencoder without the explicit construction of an order parameter. Considering that the partition function of ($d$ + 1)-dimensional anisotropic models can be mapped to that of the $d$-dimensional quantum spin models, the present study provides numerical evidence that a variational autoencoder can be applied to analyze quantum systems via the quantum Monte Carlo method.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Distributionally Robust Decentralized Volt-Var Control With Network Reconfiguration

Here, this paper presents a decentralized volt-var optimization (VVO) and network reconfiguration strategy to address the challenges arising from the growing integration of distributed energy resources, particularly photovoltaic (PV) generation units, in active distribution networks. To reconcile control measures with different time resolutions and empower local control centers to handle intermittency locally, the proposed approach leverages a two-stage distributionally robust optimization; decisions on slow-responding control measures and set points that link neighboring subnetworks are made in advance while considering all plausible distributions of uncertain PV outputs. We present a decomposition algorithm with an acceleration scheme for solving the proposed model. Numerical experiments on the IEEE 123 bus distribution system are given to demonstrate its outstanding out-of-sample performance and computational efficiency, which suggests that the proposed method can effectively localize uncertainty via risk-informed proactive timely decisions.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Fragme∩t: An Open‐Source Framework for Multiscale Quantum Chemistry Based on Fragmentation

Fragment-based quantum chemistry offers a means to circumvent the nonlinear computational scaling of conventional electronic structure calculations, by partitioning a large calculation into smaller subsystems then considering the many-body interactions between them. Variants of this approach have been used to parameterize classical force fields and machine learning potentials, applications that benefit from interoperability between quantum chemistry codes. However, there is a dearth of software that provides interoperability yet is purpose-built to handle the combinatorial complexity of fragment-based calculations. To fill this void we introduce “Fragme∩t”, an open-source software application that provides a tool for community validation of fragment-based methods, a platform for developing new approximations, and a framework for analyzing many-body interactions. Fragme∩t includes algorithms for automatic fragment generation and structure modification, and for distance- and energy-based screening of the requisite subsystems. Checkpointing, database management, and parallelization are handled internally and results are archived in a portable database. Interfaces to various quantum chemistry engines are easy to write and exist already for Q-Chem, PySCF, xTB, Orca, CP2K, MRCC, Psi4, NWChem, GAMESS, and MOPAC. Applications reported here demonstrate parallel efficiencies around 96% on more than 1000 processors but also showcase that the code can handle large-scale protein fragmentation using only workstation hardware, all with a codebase that is designed to be usable by non-experts. Fragme∩t conforms to modern software engineering best practices and is built upon well established technologies including Python, SQLite, and Ray. The source code is available under the Apache 2.0 license.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Optimized Lie–Trotter–Suzuki decompositions for two and three non-commuting terms

Lie–Trotter–Suzuki decompositions are an efficient way to approximate operator exponentials exp ( t H ) when H is a sum of n (non-commuting) terms which, individually, can be exponentiated easily. They are employed in time-evolution algorithms for tensor network states, digital quantum simulation protocols, path integral methods like quantum Monte Carlo, and splitting methods for symplectic integrators in classical Hamiltonian systems. Here, we provide optimized decompositions up to order t 6 . The leading error term is expanded in nested commutators (Hall bases) and we minimize the 1-norm of the coefficients. For n = 2 terms, several of the optima we find are close to those in McLachlan (1995). Generally, our results substantially improve over unoptimized decompositions by Forest, Ruth, Yoshida, and Suzuki. We explain why these decompositions are sufficient to efficiently simulate any one- or two-dimensional lattice model with finite-range interactions. This follows by solving a partitioning problem for the interaction graph.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗