Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “partitioned algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Ensemble transfer learning for the prediction of anti-cancer drug response

Abstract Transfer learning, which transfers patterns learned on a source dataset to a related target dataset for constructing prediction models, has been shown effective in many applications. In this paper, we investigate whether transfer learning can be used to improve the performance of anti-cancer drug response prediction models. Previous transfer learning studies for drug response prediction focused on building models to predict the response of tumor cells to a specific drug treatment. We target the more challenging task of building general prediction models that can make predictions for both new tumor cells and new drugs. Uniquely, we investigate the power of transfer learning for three drug response prediction applications including drug repurposing, precision oncology, and new drug development, through different data partition schemes in cross-validation. We extend the classic transfer learning framework through ensemble and demonstrate its general utility with three representative prediction algorithms including a gradient boosting model and two deep neural networks. The ensemble transfer learning framework is tested on benchmark in vitro drug screening datasets. The results demonstrate that our framework broadly improves the prediction performance in all three drug response prediction applications with all three prediction algorithms.

60 APPLIED LIFE SCIENCES↗

Enhancing segmentation fairness through curriculum learning and progressive loss: a centralized and federated perspective on radiograph analysis

Bias in medical image segmentation can lead to unequal performance across demographic subgroups, raising concerns about fairness and reliability in clinical AI systems. While deep learning models have achieved high segmentation accuracy, ensuring equitable performance across race and gender remains a significant challenge, particularly in privacy-sensitive healthcare environments. This study investigates fairness-aware medical image segmentation for hip and knee radiographs using deep learning models evaluated in both centralized and Federated Learning (FL) settings. We introduce Curriculum Learning (CL) strategies and Progressive Loss (PL) functions to regulate sample difficulty during training. In addition, we propose two novel fairness-oriented federated learning algorithms, Federated Intersection over Union (FedIoU) and Federated Intersection over Union with Outlier Analysis (FedIoUoutlier). Experiments are conducted using multiple segmentation backbones and simulated multi-site data partitions derived from the Osteoarthritis Initiative dataset. Model performance is evaluated using Intersection over Union (IoU), IoU standard deviation, Skewed Error Ratio (SER), and Min-Max Disparity across race and gender subgroups. Statistical significance was verified using paired t-tests to compare per-sample IoU performance against baseline configurations. Across both hip and knee segmentation tasks, curriculum learning and progressive loss strategies consistently improved segmentation accuracy and reduced demographic performance disparities in centralized training. In federated settings, fairness-aware aggregation further enhanced performance. Notably, FedIoUoutlier combined with balanced curriculum learning and tiered progressive loss achieved the highest mean IoU while yielding the lowest SER and Min-Max Disparity, indicating improved fairness without sacrificing accuracy. In several configurations, federated models matched or exceeded the performance of optimized centralized models, with statistically significant improvements in per-sample IoU over baseline configurations. The results demonstrate that structured training strategies and fairness-aware federated aggregation can jointly improve accuracy, stability, and demographic fairness in medical image segmentation. By integrating curriculum learning, progressive loss, and novel FL algorithms, this work provides a practical pathway toward equitable and privacy-preserving AI systems for medical imaging.

97 MATHEMATICS AND COMPUTING↗

GronOR: Massively Parallel and GPU-Accelerated Non-Orthogonal Configuration Interaction for Large Molecular Systems

GronOR is a program package for non-orthogonal configuration interaction calculations for an electronic wave function built in terms of anti-symmetrized products of multi-configuration molecular fragment wave functions. The two-electron integrals that have to be processed may be expressed in terms of atomic orbitals or in terms of an orbital basis determined from the molecular orbitals of the fragments. The code has been specifically designed for execution on distributed memory massively parallel and Graphics Processing Unit (GPU)-accelerated computer architectures, using an MPI+OpenACC/OpenMP programming approach. The task-based execution model used in the implementation allows for linear scaling with the number of nodes on the largest pre-exascale architectures available, provides hardware fault resiliency, and enables effective execution on systems with distinct central processing unit-only and GPU-accelerated partitions. The code interfaces with existing multi-configuration electronic structure codes that provide optimized molecular fragment orbitals, configuration interaction coefficients, and the required integrals. Algorithm and implementation details, parallel and accelerated performance benchmarks, and an analysis of the sensitivity of the accuracy of results and computational performance to thresholds used in the calculations are presented.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A distributed voltage inference framework for cyber-physical attacks detection and localization in active distribution grids

The transition to active distribution grids with real-time monitoring and control depends on the proliferation of advanced communication networks and devices. This paradigm shift towards a cyber-physical architecture also introduces new vulnerabilities for adversaries to exploit and launch sophisticated cyber-physical attacks targeting grid observability. Current research highlights the challenges in distinguishing attacks on voltage phasor or nodal injection measurements and isolating multi-source attack locations in a multiphase distribution grid. The attack detection and localization methods in literature face accuracy issues, applications across diverse attack scenarios, or scalability limits. Here, to bridge these gaps, this paper proposes a distributed Voltage Inference framework for real-time detection and localization of cyber-physical attacks, addressing scalability, adaptability, and accuracy challenges in state-of-the-art methods. The proposed methodology leverages the distributed nature of the Voltage Inference framework through a two-step process of prediction and correction, together with a tractable graph partitioning approach, providing a reliable solution to identify compromised measurement sources and facilitate isolation. Extensive testing on IEEE 13 and 123-node distribution feeders underscores the algorithm’s efficacy, enhancing the security and resilience of active distribution grids against evolving cyber threats. Additionally, Hardware-in-the-Loop (HIL) implementation validates the proposed strategy’s practical applicability in real-world scenarios.

active distribution grids↗

A Stabilizer Framework for the Contextual Subspace Variational Quantum Eigensolver and the Noncontextual Projection Ansatz

Quantum chemistry is a promising application for noisy intermediate-scale quantum (NISQ) devices. However, quantum computers have thus far not succeeded in providing solutions to problems of real scientific significance, with algorithmic advances being necessary to fully utilize even the modest NISQ machines available today. We discuss a method of ground state energy estimation predicated on a partitioning of the molecular Hamiltonian into two parts: one that is noncontextual and can be solved classically, supplemented by a contextual component that yields quantum corrections obtained via a Variational Quantum Eigensolver (VQE) routine. This approach has been termed Contextual Subspace VQE (CS-VQE); however, there are obstacles to overcome before it can be deployed on NISQ devices. The problem we address here is that of the ansatz, a parametrized quantum state over which we optimize during VQE; it is not initially clear how a splitting of the Hamiltonian should be reflected in the CS-VQE ansätze. We propose a “noncontextual projection” approach that is illuminated by a reformulation of CS-VQE in the stabilizer formalism. This defines an ansatz restriction from the full electronic structure problem to the contextual subspace and facilitates an implementation of CS-VQE that may be deployed on NISQ devices. We validate the noncontextual projection ansatz using a quantum simulator and demonstrate chemically precise ground state energy calculations for a suite of small molecules at a significant reduction in the required qubit count and circuit depth.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Implementing Directive-Based Deferred Execution for Effective Network Aggregation

Remote direct memory access technology provides an efficient mechanism for one-sided communication that can be leveraged to implement a distributed shared memory programming model. However, when applications generate large numbers of small, irregular messages, network congestion often arises. Existing solutions address this small message problem by facilitating message aggregation but typically require disruptive code transformations that detract from the algorithmic intent of applications, or can be limited by dependent operations on aggregated data between synchronisation points. A solution is to use a directive-assisted approach that enables compilers to transform code dependent on aggregated communication for deferred execution. This paper presents an algorithm that a compiler can use to implement and optimise deferred execution for code dependent on aggregated data, based on an "aggregation context" extension for the OpenSHMEM partitioned global address space library. This new capability addresses a key challenge of message aggregation, allowing its full potential to reduce network congestion and enhance programmability to be realised.

Welch, Aaron [ORNL]↗

Adaptive Hierarchical Cyber Attack Detection and Localization in Active Distribution Systems

Development of a cyber security strategy for the active distribution systems is challenging due to the inclusion of distributed renewable energy generations. Here this paper proposes an adaptive hierarchical cyber attack detection and localization framework for distributed active distribution systems via analyzing electrical waveforms. Cyber attack detection is based on a sequential deep learning model, via which even minor cyber attacks can be identified. The two-stage cyber attack localization algorithm first estimates the cyber attack sub-region, and then localize the specified cyber attack within the estimated subregion. We propose a modified spectral clustering-based network partitioning method for the hierarchical cyber attack ‘coarse’ localization. Next, to further narrow down the cyber attack location, a normalized impact score based on waveform statistical metrics is proposed to obtain a ‘fine’ cyber attack location by characterizing different waveform properties. Finally, compared with classical and state-of-art methods, a comprehensive quantitative evaluation with two case studies shows promising estimation results of the proposed framework.

42 ENGINEERING↗

Predicting secondary organic aerosol phase state and viscosity and its effect on multiphase chemistry in a regional-scale air quality model

Atmospheric aerosols are a significant public health hazard and have substantial impacts on the climate. Secondary organic aerosols (SOAs) have been shown to phase separate into a highly viscous organic outer layer surrounding an aqueous core. This phase separation can decrease the partitioning of semi-volatile and low-volatile species to the organic phase and alter the extent of acid-catalyzed reactions in the aqueous core. A new algorithm that can determine SOA phase separation based on their glass transition temperature (T g ), oxygen to carbon (O:C) ratio and organic mass to sulfate ratio, and meteorological conditions was implemented into the Community Multiscale Air Quality Modeling (CMAQ) system version 5.2.1 and was used to simulate the conditions in the continental United States for the summer of 2013. SOA formed at the ground/surface level was predicted to be phase separated with core–shell morphology, i.e., aqueous inorganic core surrounded by organic coating 65.4 % of the time during the 2013 Southern Oxidant and Aerosol Study (SOAS) on average in the isoprene-rich southeastern United States. Our estimate is in proximity to the previously reported ~70 % in literature. The phase states of organic coatings switched between semi-solid and liquid states, depending on the environmental conditions. The semi-solid shell occurring with lower aerosol liquid water content (western United States and at higher altitudes) has a viscosity that was predicted to be 10 2 –10 12 Pa s, which resulted in organic mass being decreased due to diffusion limitation. Organic aerosol was primarily liquid where aerosol liquid water was dominant (eastern United States and at the surface), with a viscosity <10 2 Pa s. Phase separation while in a liquid phase state, i.e., liquid–liquid phase separation (LLPS), also reduces reactive uptake rates relative to homogeneous internally mixed liquid morphology but was lower than aerosols with a thick viscous organic shell. The sensitivity cases performed with different phase-separation parameterization and dissolution rate of isoprene epoxydiol (IEPOX) into the particle phase in CMAQ can have varying impact on fine particulate matter (PM 2.5 ) organic mass, in terms of bias and error compared to field data collected during the 2013 SOAS. This highlights the need to better constrain the parameters that govern phase state and morphology of SOA, as well as expand mechanistic representation of multiphase chemistry for non-IEPOX SOA formation in models aided by novel experimental insights.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Recent Advances of PyROS: A Pyomo Solver for Nonconvex Two-Stage Robust Optimization in Process Systems Engineering

The document presents recent algorithmic and implementation advances of the two-stage robust optimization (RO) solver PyROS, and a benchmarking study which demonstrates the utility of PyROS for two-stage RO problems. The advances include extensions of the scope of PyROS to models with uncertain variable bounds, improvements to the initializations of the subproblems used by the underlying cutting set algorithm, and extensions of the uncertainty set interfaces. The benchmarking study is performed on a library of over 8,500 instances, with variations in the nonlinearities, degree-of-freedom partitioning, uncertainty sets, and polynomial decision rule approximations. An amine-based CO2 capture case study is presented to demonstrate the utility of PyROS for large-scale process models. Overall, the results highlight the effectiveness of PyROS for obtaining robust solutions to optimization problems with uncertain equality constraints.

Sherman, Jason↗

Recent Advances in PyROS: The Pyomo Solver for Two-Stage Nonconvex Robust Optimization

The slides present recent algorithmic and implementation advances of the two-stage robust optimization (RO) solver PyROS, and a benchmarking study which demonstrates the utility of PyROS for two-stage RO problems. The advances include extensions of the scope of PyROS to models with uncertain variable bounds, improvements to the initializations of the subproblems used by the underlying cutting set algorithm, and extensions of the uncertainty set interfaces. The benchmarking study is performed on a library of over 8,500 instances, with variations in the nonlinearities, degree-of-freedom partitioning, uncertainty sets, and polynomial decision rule approximations. Overall, the results highlight the effectiveness of PyROS for obtaining robust solutions to optimization problems with uncertain equality constraints.

Sherman, Jason↗

Recent Advances in PyROS: The Pyomo Solver for Two-Stage Nonconvex Robust Optimization

The slides present recent algorithmic and implementation advances of the two-stage robust optimization (RO) solver PyROS, and a benchmarking study which demonstrates the utility of PyROS for two-stage RO problems. The advances include extensions of the scope of PyROS to models with uncertain variable bounds, improvements to the initializations of the subproblems used by the underlying cutting set algorithm, and extensions of the uncertainty set interfaces. The benchmarking study is performed on a library of over 8,500 instances, with variations in the nonlinearities, degree-of-freedom partitioning, uncertainty sets, and polynomial decision rule approximations. Overall, the results highlight the effectiveness of PyROS for obtaining robust solutions to optimization problems with uncertain equality constraints.

Sherman, Jason↗

Enabling Efficient Sparse Computations using Linear Algebra Aware Compilers

This project developed the LAPIS compiler framework, built on the Multilevel Intermediate Representation (MLIR), to optimize sparse linear algebra operations and support performance portability across diverse architectures. The main innovation of LAPIS is the Kokkos dialect, which allows for lowering codes from a high productivity language to different architectures in an elegant way. The dialect also allows the conversion of lower-level MLIR code to C++ Kokkos code, facilitating the integration of scientific machine learning (SciML) models into applications. To extend LAPIS for distributed memory architectures, a new partition dialect was created to manage the distribution of sparse tensors and express communication patterns for sparse linear algebra operations. This dialect also supports the distributed execution of operators and includes algorithmic optimizations to minimize communication to improve performance. The project also demonstrates that MLIR can enable effective linear algebra-level optimizations, improving performance on different GPUs for both sparse and dense linear algebra kernels. Key applications of LAPIS include sparse linear algebra and graph kernels, TenSQL, a relational database management solution built on GraphBLAS, and the development of subgraph isomorphism and monomorphism kernels, showcasing performance portability. In summary, the LAPIS framework supports productivity, performance, portability, and distributed memory execution, while also enabling linear algebra-level optimizations that are challenging in traditional programming languages, with successful applications ranging from simple sparse linear algebra to complex graph kernels.

97 MATHEMATICS AND COMPUTING↗

Learning linear optical circuits with coherent states

We analyze the energy and training data requirements for supervised learning of an M-mode linear optical circuit by minimizing an empirical risk defined solely from the action of the circuit on coherent states. When the linear optical circuit acts non-trivially only on k < M unknown modes (i.e. a linear optical k-junta), we provide an energy-efficient, adaptive algorithm that identifies the junta set and learns the circuit. We compare two schemes for allocating a total energy, E, to the learning algorithm. In the first scheme, each of the T random training coherent states has energy E/T. In the second scheme, a single random MT-mode coherent state with energy E is partitioned into T training coherent states. The latter scheme exhibits a polynomial advantage in training data size sufficient for convergence of the empirical risk to the full risk due to concentration of measure on the $(2MT-1)$-sphere. Specifically, generalization bounds for both schemes are proven, which indicate that for ε-approximation of the full risk by the empirical risk with high probability, $O(E^{2/3}M^{2/3}/\epsilon^{2/3})$ training states are sufficient for the first scheme and $O(E^{1/3}M^{1/3}/\epsilon^{2/3})$ training states are sufficient for the second scheme.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Robust Multi-fidelity Bayesian Optimization with Deep Kernel and Partition

Multi-fidelity Bayesian optimization (MFBO) is a powerful approach that utilizes lowfidelity, cost-effective sources to expedite the exploration and exploitation of a high-fidelity objective function. Existing MFBO methods with theoretical foundations either lack justification for performance improvements over single-fidelity optimization or rely on strong assumptions about the relationships between fidelity sources to construct surrogate models and direct queries to low-fidelity sources. To mitigate the dependency on cross-fidelity assumptions while maintaining the advantages of low-fidelity queries, we introduce a random sampling and partition-based MFBO framework with deep kernel learning. This framework is robust to cross-fidelity model misspecification and explicitly illustrates the benefits of low-fidelity queries. Our results demonstrate that the proposed algorithm effectively manages complex cross-fidelity relationships and efficiently optimizes the target fidelity function.

Zhang, Fengxue [University of Chicago, Illinois, U↗

Profiling the BLAST bioinformatics application for load balancing on high-performance computing clusters

Abstract Background The Basic Local Alignment Search Tool (BLAST) is a suite of commonly used algorithms for identifying matches between biological sequences. The user supplies a database file and query file of sequences for BLAST to find identical sequences between the two. The typical millions of database and query sequences make BLAST computationally challenging but also well suited for parallelization on high-performance computing clusters. The efficacy of parallelization depends on the data partitioning, where the optimal data partitioning relies on an accurate performance model. In previous studies, a BLAST job was sped up by 27 times by partitioning the database and query among thousands of processor nodes. However, the optimality of the partitioning method was not studied. Unlike BLAST performance models proposed in the literature that usually have problem size and hardware configuration as the only variables, the execution time of a BLAST job is a function of database size, query size, and hardware capability. In this work, the nucleotide BLAST application BLASTN was profiled using three methods: shell-level profiling with the Unix “time” command, code-level profiling with the built-in “profiler” module, and system-level profiling with the Unix “gprof” program. The runtimes were measured for six node types, using six different database files and 15 query files, on a heterogeneous HPC cluster with 500+ nodes. The empirical measurement data were fitted with quadratic functions to develop performance models that were used to guide the data parallelization for BLASTN jobs. Results Profiling results showed that BLASTN contains more than 34,500 different functions, but a single function, RunMTBySplitDB, takes 99.12% of the total runtime. Among its 53 child functions, five core functions were identified to make up 92.12% of the overall BLASTN runtime. Based on the performance models, static load balancing algorithms can be applied to the BLASTN input data to minimize the runtime of the longest job on an HPC cluster. Four test cases being run on homogeneous and heterogeneous clusters were tested. Experiment results showed that the runtime can be reduced by 81% on a homogeneous cluster and by 20% on a heterogeneous cluster by re-distributing the workload. Discussion Optimal data partitioning can improve BLASTN’s overall runtime 5.4-fold in comparison with dividing the database and query into the same number of fragments. The proposed methodology can be used in the other applications in the BLAST+ suite or any other application as long as source code is available.

59 BASIC BIOLOGICAL SCIENCES↗

Jacobian-scaled K-means clustering for physics-informed segmentation of reacting flows

This work introduces Jacobian-scaled K-means (JSK-means) clustering, which is a physicsinformed clustering strategy centered on the K-means framework. The method allows for the injection of underlying physical knowledge into the clustering procedure through a distance function modification: instead of leveraging conventional Euclidean distance vectors, the JSKmeans procedure operates on distance vectors scaled by matrices obtained from dynamical system Jacobians evaluated at the cluster centroids. The goal of this work is to show how the JSKmeans algorithm - without modifying the input dataset - produces clusters that capture regions of dynamical similarity, in that the clusters are redistributed towards high-sensitivity regions in phase space and are described by similarity in the source terms of samples instead of the samples themselves. The algorithm is demonstrated on a complex reacting flow simulation dataset (a channel detonation configuration), where the dynamics in the thermochemical composition space are known through the highly nonlinear and stiff Arrhenius-based chemical source terms. Interpretations of cluster partitions in both physical space and composition space reveal how JSK-means shifts clusters produced by standard K-means towards regions of high chemical sensitivity (e.g., towards regions of peak heat release rate near the detonation reaction zone). Furthermore, the findings presented here illustrate the benefits of utilizing Jacobian-scaled distances in clustering techniques, and the JSK-means method in particular displays promising potential for improving former partition-based modeling strategies in reacting flow (and other multi-physics) applications.

Clustering↗

Systemwide Planning with a Branch-and-Price Algorithm for Pavement-Marking Assessment Data Collection via the Mobile Retroreflectivity Unit Routing Model

The visibility of pavement markings is one of the most critical factors for traffic safety, and a periodical assessment plan is crucial for maintaining this function. Traditional assessment methods, such as visual windshield surveys or manual testing using handheld devices, are unsafe, time-consuming, and labor-intensive. In recent years, transportation agencies have begun to adopt the use of mobile retroreflectivity units (MRUs) for condition assessment of pavement markings. MRUs, different from other manual methods, can be utilized to collect large-scale retroreflectivity data in an efficient manner. However, no relevant research has yet proposed a mathematical optimization model for arranging the evaluation schedule and paths of MRUs. This study aims to propose a MRU routing model, and an efficient solution methodology. A branch-and-price algorithm, including column generation and branch-and-bound, was implemented. Computational experiments have been conducted based on actual tasks from the Florida MRU program for validation. In conclusion, results show that the proposed solution methodology with a set partitioning model in this study not only finds the optimal solution for problems with tasks less than 60, but also effectively narrows the solution gap to be within 1.0% for problems with tasks less than 131.

42 ENGINEERING↗

Deep Reinforcement Learning Enabled Physical-Model-Free Two-Timescale Voltage Control Method for Active Distribution Systems

Active distribution networks are being challenged by frequent and rapid voltage violations due to renewable energy integration. Conventional model-based voltage control methods rely on accurate parameters of the distribution networks, which are difficult to achieve in practice. This paper proposes a novel physical-model-free two-timescale voltage control framework for active distribution systems. To achieve fast control of PV inverters, the whole network is first partitioned into several subnetworks using voltage-reactive power sensitivity. Then, the scheduling of PV inverters in the multiple sub-networks is formulated as Markov games and solved by a multi-agent soft actor-critic (MASAC) algorithm, where each subnetwork is modeled as an intelligent agent. All agents are trained in a centralized manner to learn a coordinated strategy while being executed based on only local information for fast response. For the slower time-scale control, OLTCs and switched capacitors are coordinated by a single agent-based SAC algorithm using the global information with considering control behaviors of the inverters. Particularly, the two-level agents are trained concurrently with information exchange according to the reward signal calculated from the data-driven surrogate model. Comparative tests with different benchmark methods on IEEE 33-and 123-bus systems and 342-node low voltage distribution system demonstrate that the proposed method can effectively mitigate the fast voltage violations and achieve systematical coordination of different voltage regulation assets without the knowledge of accurate system model.

24 POWER TRANSMISSION AND DISTRIBUTION↗