Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Massively parallel and universal approximation of nonlinear functions using diffractive processors

Nonlinear computation is essential for a wide range of information processing tasks, yet implementing nonlinear functions using optical systems remains a challenge due to the weak and power-intensive nature of optical nonlinearities. Overcoming this limitation without relying on nonlinear optical materials could unlock unprecedented opportunities for ultrafast and parallel optical computing systems. Here, we demonstrate that large-scale nonlinear computation can be performed using linear optics through optimized diffractive processors composed of passive phase-only surfaces. In this framework, the input variables of nonlinear functions are encoded into the phase of an optical wavefront—e.g., via a spatial light modulator (SLM)—and transformed by an optimized diffractive structure with spatially varying point-spread functions to yield output intensities that approximate a large set of unique nonlinear functions–all in parallel. We provide proof establishing that this architecture serves as a universal function approximator for an arbitrary set of bandlimited nonlinear functions, also covering wavelength-multiplexed nonlinear functions as well as multi-variate and complex-valued functions that are all-optically cascadable. Our analysis also indicates the successful approximation of typical nonlinear activation functions commonly used in neural networks, including the sigmoid, tanh, ReLU (rectified linear unit), and softplus. We numerically demonstrate the parallel computation of one million distinct nonlinear functions, accurately executed at wavelength-scale spatial density at the output of a diffractive optical processor. Furthermore, we experimentally validated this framework using in situ optical learning and approximated 35 unique nonlinear functions in a single shot using a compact setup consisting of an SLM and an image sensor. These results establish diffractive optical processors as a scalable platform for massively parallel universal nonlinear function approximation, paving the way for new capabilities in analog optical computing based on linear materials.

Rahman, Md Sadman Sakib [University of California,↗

Mitigating Cascading Outages in Severe Weather Using Simulation-Based Optimization

Severe weather events can trigger cascading power outages and lead to significant losses. In this work, we investigate cascading outage mitigation under severe weather conditions. Given day-ahead weather forecasts and component failure models, we aim to identify a set of power lines that can be hardened to minimize the expected impact of potential cascading outages. Since the expected load shedding cannot be expressed as an explicit function of line hardening decisions and system states, we developed a cascading outage simulator to estimate the expected value of load shedding under various initial weather-related disruption scenarios generated using a weather forecast. To avoid massive enumeration of all possible combinations of line hardening decisions and reduce the simulation efforts, we employed an efficient simulation-based optimization approach that quickly identifies the (near) optimal line hardening decisions in the presence of both large simulation noises due to the highly variable initial disturbances and system states, and significant randomness in the subsequent cascades. Furthermore, the algorithm is also able to utilize parallel computing to dramatically reduce computation time to support decision making in preparation for severe weather conditions. We performed a case study on the Northeast Power Coordinating Council (NPCC) 140-bus system model to demonstrate that our approach can significantly improve power grid resilience to adverse weather events.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Source-to-Source Automatic Differentiation of OpenMP Parallel Loops

This article presents our work toward correct and efficient automatic differentiation of OpenMP parallel worksharing loops in forward and reverse mode. Automatic differentiation is a method to obtain gradients of numerical programs, which are crucial in optimization, uncertainty quantification, and machine learning. The computational cost to compute gradients is a common bottleneck in practice. For applications that are parallelized for multicore CPUs or GPUs using OpenMP, one also wishes to compute the gradients in parallel. Here, we propose a framework to reason about the correctness of the generated derivative code, from which we justify our OpenMP extension to the differentiation model. We implement this model in the automatic differentiation tool Tapenade and present test cases that are differentiated following our extended differentiation procedure. Performance of the generated derivative programs in forward and reverse mode is better than sequential, although our reverse mode often scales worse than the input programs.

97 MATHEMATICS AND COMPUTING↗

ParChain: a framework for parallel hierarchical agglomerative clustering using nearest-neighbor chain

This paper studies the hierarchical clustering problem, where the goal is to produce a dendrogram that represents clusters at varying scales of a data set. We propose the ParChain framework for designing parallel hierarchical agglomerative clustering (HAC) algorithms, and using the framework we obtain novel parallel algorithms for the complete linkage, average linkage, and Ward's linkage criteria. Compared to most previous parallel HAC algorithms, which require quadratic memory, our new algorithms require only linear memory, and are scalable to large data sets. ParChain is based on our parallelization of the nearest-neighbor chain algorithm, and enables multiple clusters to be merged on every round. We introduce two key optimizations that are critical for efficiency: a range query optimization that reduces the number of distance computations required when finding nearest neighbors of clusters, and a caching optimization that stores a subset of previously computed distances, which are likely to be reused. Experimentally, we show that our highly-optimized implementations using 48 cores with two-way hyper-threading achieve 5.8--110.1x speedup over state-of-the-art parallel HAC algorithms and achieve 13.75--54.23x self-relative speedup. Compared to state-of-the-art algorithms, our algorithms require up to 237.3x less space. Our algorithms are able to scale to data set sizes with tens of millions of points, which existing algorithms are not able to handle.

Computer Science↗

Parallel Battery: The Framework and Process for an Intelligent and Ecological Battery System and Related Services

The concept, framework, process methodology and applications of parallel battery were proposed from both virtual and real aspects.The parallel battery was an application of ACP-based parallel intelligence in battery and related energy system areas.The real battery system was running with its equivalent, and the artificial battery system was in a virtual space, in a parallel and interactive manner.The artificial battery system contained the descriptive, predictive, and prescriptive functions on the real battery and related energy systems.There was a closed-loop workflow between the real battery system and the artificial battery system, which iteratively optimizes the battery and related energy systems, leading to a new paradigm of intelligent and ecological parallel battery system management.

ACP approach↗

Dynamic Ride-Matching for Large-Scale Transportation Systems

Efficient dynamic ride-matching (DRM) in large-scale transportation systems is a key driver in transport simulations to yield answers to challenging problems. Although the DRM problem is simple to solve, it quickly becomes a computationally challenging problem in large-scale transportation system simulations. Therefore, this study thoroughly examines the DRM problem dynamics and proposes an optimization-based solution framework to solve the problem efficiently. To benefit from parallel computing and reduce computational times, the problem’s network is divided into clusters utilizing a commonly used unsupervised machine learning algorithm along with a linear programming model. Then, these sub-problems are solved using another linear program to finalize the ride-matching. At the clustering level, the framework allows users adjusting cluster sizes to balance the trade-off between the computational time savings and the solution quality deviation. A case study in the Chicago Metropolitan Area, U.S., illustrates that the framework can reduce the average computational time by 58% at the cost of increasing the average pick up time by 26% compared with a system optimum, that is, non-clustered, approach. Another case study in a relatively small city, Bloomington, Illinois, U.S., shows that the framework provides quite similar results to the system-optimum approach in approximately 62% less computational time.

33 ADVANCED PROPULSION SYSTEMS↗

Stability Analysis of Parallel Connected Bidirectional WPT System

This paper presents a stability analysis of parallel-connected bi-directional series-series resonant network wireless power transfer (WPT), optimized for Electric Vehicle (EV) charging and vehicle-to-grid (V2G) applications. The study addresses critical stability challenges in systems integrated with diverse distributed energy resources (DERs), including photovoltaics, fuel cells, wind turbines, energy storage systems, and the AC grid. The stability of such integrated DC grid systems is paramount for ensuring reliable operation, particularly under varying power flow conditions and dynamic interactions between parallel WPT systems. The analysis included system impedance characterization, state-space modeling, and open and closed-loop stability evaluations. The results demonstrated that the integration of a robust control architecture effectively mitigates instability risks and supports scalable, efficient operation. This work underscores the converter's adaptability and its potential for large-scale deployment in wireless EV charging infrastructures and integrated DC grid systems.

Asa, Erdem [ORNL] (ORCID:0000000190884812)↗

Accelerating Simulation for High-Fidelity PV Inverter System Reliability Assessment with High-Performance Computing

The overall cost of photovoltaic (PV) systems has shown a downward trend during the last decade; however, PV inverter failures account for the highest cost of operation and maintenance. To address this, reliability tools with powerful computation and better accuracy are required for the lifetime prediction and degradation evaluation of PV inverters. This paper proposes an event-driven parallel computing-based simulator. The proposed simulator applies high-performance computing techniques and other accessory optimization techniques-including cluster merging, adaptive model updates, and steady-state identification-to make reliability assessments for PV inverters under given input mission profiles and operating conditions with high efficiency and high fidelity. The main idea of the simulator and its workflow are introduced. Then, a demo PV inverter system simulator is implemented, and the speedup of the total simulations of the switching model reaches 123.03 times.

high-performance computing↗

MOOSE ProbML: Parallelizable Probabilistic Machine Learning and Uncertainty Quantification Capabilities

The Multiphysics Object Oriented Simulation Environment (MOOSE) is a widely used open- source finite element software for performing multiphysics multiscale simulations in a massively parallel fashion. Recently, the computational team at Idaho National Laboratory (INL) has implemented Probabilistic Machine Learning (ProbML) capabilities in MOOSE—in a parallelized fashion—and enable active learning with large-scale computational models for tasks such as surrogate model development, scale bridging, forward/inverse uncertainty quantification (UQ), Bayesian optimization, etc. This presentation summarizes these developments in MOOSE along with demonstrations on several real applications relevant to nuclear energy. At the fundamental level, samplers like Monte Carlo/Latin Hypercube, variance reduction, parallelized Markov Chain Monte Carlo (MCMC) support uncertainty propagation in both forward and inverse settings. These samplers can be integrated with the Gaussian processes (GP) suite in MOOSE, which offer several variants like scalar GPs, multi-output GPs, and deep GPs, to enable active learning. These GPs can be tuned using gradient-based optimization methods like Adam and its variants or gradient-free methods like the elliptical slice sampler (a variant of MCMC adept under Gaussian settings) for more complex covariance kernels or likelihoods whose gradient computations can be cumbersome. A variety of batch acquisition functions permit parallelized evaluation of the computational model and support different learning objectives with high efficiency like Bayesian inference, global surrogate development, optimization, etc. Furthermore, libtorch integration supports training, evaluation, and re-training of neural networks and other complex machine learning models in active learning settings. The impacts of these developments are shown on several real applications: (1) nuclear fuel inverse UQ and model inadequacy assessment using the Kennedy O’Hagan framework; (2) uncertainty aware surrogate modeling for additive manufacturing to predict field quantities; (3) nuclear reactor rare events analysis; and (4) complex fluid flow prediction using a global surrogate with quantified prediction uncertainty. Finally, the outlook of MOOSE ProbML is discussed for both outer-loop and inner-loop computations in the broad view to accelerate fuels and materials qualification, address gaps in knowledge and data, and assess new reactor/fuel systems.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Efficient Parallelization of Irregular Applications on GPU Architectures

With the enlarging computation capacity of general Graphics Processing Units (GPUs), leveraging GPUs to accelerate parallel applications has become a critical topic in academia and industry. However, a wide range of irregular applications with the computation-/memory-intensive nature cannot easily achieve high GPU utilization. The challenges mainly involve the following aspects: first, data dependence leads to coarse-grained kernel and inefficient parallelism; second, heavy GPU memory usage may cause frequent memory evictions and extra overhead of I/O; third, specific computation patterns produce memory redundancies; last, workload balance and data reusability conjunctly benefit the overall performance, but there may exist a dynamic trade-off between them. Targeting these challenges, this dissertation proposes multiple optimizations to accelerate two real-world applications: many-body correlation functions to simulate nuclear physics in a large-scale scientific system; the other is the eALS-based matrix factorization recommendation system. To accelerate the calculations of many-body correlation functions, this dissertation presents three frameworks in GPU memory management and multi-GPU scheduling. Firstly, an optimized systematic GPU memory management framework, MemHC, utilizes a series of new memory reduction designs in GPU memory allocation, CPU/GPU communications, and GPU memory oversubscription. Secondly, an enhanced multi-GPU scheduling framework, MICCO, particularly by taking both data dimension (e.g., data reuse and data eviction) and computation dimension into account. MICCO designs a heuristic scheduling algorithm and a machine learning-based regression model to generate the optimal settings of a proposed new concept to manage the trade-off. Thirdly, a locality-aware multi-GPU scheduling framework. This scheduler leverages pipeline batch generation with a looking-ahead strategy by building local dependency graphs for memory transfer reduction and better data reuse, achieving up to 79.92% memory cost reduction and 1.67x speedup. To parallelize the eALS-based recommendation system, this dissertation proposes an efficient CPU/GPU heterogeneous recommendation system, HEALS. HEALS employs newly designed architecture-adaptive data formats to achieve load balance and good data locality on CPU and GPU. To mitigate the data dependence, HEALS presents a CPU/GPU collaboration model for both task parallelism and data parallelism with multiple kernel computation optimizations. In summary, this dissertation efficiently accelerates two typical irregular applications on GPUs by building four frameworks, including CPU/GPU collaboration, GPU memory management, and multi-GPU scheduling.

Wang, Qihan↗

Active Learning for Metamaterial Optimization on HPC and QC Integrated Systems

Active learning algorithms, integrating machine learning, quantum computing and optics simulation in an iterative loop, offer a promising approach to optimizing metamaterials. However, these algorithms can face difficulties in optimizing highly complex structures due to computational limitations. High-performance computing (HPC) and quantum computing (QC) integrated systems can address these issues by enabling parallel computing. In this study, we develop an active learning algorithm working on HPC-QC integrated systems. We evaluate the performance of optimization processes within active learning (i.e., training a machine learning model, problem-solving with quantum computing, and evaluating optical properties through wave-optics simulation) for highly complex metamaterial cases. Our results showcase that utilizing multiple cores on the integrated system can significantly reduce computational time, thereby enhancing the efficiency of optimization processes. Therefore, we expect that leveraging HPC-QC integrated systems helps effectively tackle large-scale optimization challenges in general.

Kim, Seongmin↗

Asynchronous Richardson iterations: theory and practice

We consider asynchronous versions of the first- and second-order Richardson methods for solving linear systems of equations. These methods depend on parameters whose values are chosen a priori. We explore the parameter values that can be proven to give convergence of the asynchronous methods. This is the first such analysis for asynchronous second-order methods. We find that for the first-order method, the optimal parameter value for the synchronous case also gives an asynchronously convergent method. For the second-order method, the parameter ranges for which we can prove asynchronous convergence do not contain the optimal parameter values for the synchronous iteration. In practice, however, the asynchronous second-order iterations may still converge using the optimal parameter values, or parameter values close to the optimal ones, despite this result. We explore this behavior with a multithreaded parallel implementation of the asynchronous methods.

97 MATHEMATICS AND COMPUTING↗

Development of a Reactive Force Field for Simulating Photoinitiated Acrylate Polymerization

Light-driven and photo-curable polymer based additive manufacturing (AM) has enormous potential due to its excellent resolution and precision. Acrylated radical chain-growth polymerized resins are widely used in photopolymer AM due to their fast kinetics, and often serve as a departure point for developing other resin materials for photopolymer-based AM technologies. For successful control of the photopolymer resins, the molecular basis of the acrylate free-radical polymerization has to be understood in detail. We present an optimized reactive force field (ReaxFF) for molecular dynamics (MD) simulations of acrylate polymer resins that captures radical polymerization thermodynamics and kinetics. The force field is trained against an extensive training set including density functional theory (DFT) calculations of reaction pathways along the radical polymerization from methyl acrylate to methyl butyrate, bond dissociation energies, and structures and partial charges of several molecules and radicals. We also found that it was critical to train the force field against an incorrect, nonphysical reaction pathway observed in simulations that used parameters not optimized for acrylate polymerization. As a result, the parameterization process utilizes a parallelized search algorithm, and the resulting model can describe polymer resin formation, crosslinking density, conversion rate, and residual monomers of the complex acrylate mixtures.

36 MATERIALS SCIENCE↗

Spatially Resolved Mapping of Three-Dimensional Molecular Orientations with ~2 nm Spatial Resolution through Tip-Enhanced Raman Scattering

We record local optical field images of silver nanocubes (75 nm) using tip-enhanced Raman (TER) spectral imaging. The images that we observe are consistent with several recent reports from our group, but here, we demonstrate sub-2 nm spatial resolution in local optical field nanoimaging under ambient laboratory conditions. This is achieved by scanning the substrate (nanocube on Si) relative to a 4-thiobenzonitrile (TBN)-functionalized Ag-coated TER probe. The spatial resolution that we obtain necessitates that only a few molecules govern the recorded optical response; molecular orientation becomes an important consideration in such measurements. We model the orientation through geometry optimization of a TBN molecule chemisorbed onto an Ag79 cluster (sphere with a ~1 nm diameter). Using the computed orientation of the cluster-bound molecule, we then model the optical response using formalism that accounts for the orientation of the molecule relative to vector components of the local optical fields. We find optimal agreement between experiment and theory. In effect, this work reveals the parallels between single-molecule Raman scattering and high-spatial-resolution TER spectroscopy, even when the images themselves cannot be used to visualize a single molecule in real space.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Exaflops Biomedical Knowledge Graph Analytics

We are motivated by newly proposed methods for mining large-scale corpora of scholarly publications (e.g., full biomedical literature), which consists of tens of millions of papers spanning decades of research. In this setting, analysts seek to discover relationships among concepts. They construct graph representations from annotated text databases and then formulate the relationship-mining problem as an all-pairs shortest paths (APSP) and validate connective paths against curated biomedical knowledge graphs (e.g., Spoke). In this context, we present Coast (Exascale Communication-Optimized All-Pairs Shortest Path) and demonstrate 1.004 EF/s on 9,200 Frontier nodes (73,600 GCDs). We develop hyperbolic performance models (HYPERMOD), which guide optimizations and parametric tuning. The proposed Coast algorithm achieved the memory constant parallel efficiency of 99% in the single-precision tropical semiring. Looking forward, Coast will enable the integration of scholarly corpora like PubMed into the Spoke biomedical knowledge graph.

Kannan, Ramakrishnan {ramki}↗

SPEL: Software tool for Porting E3SM Land Model with OpenACC in a Function Unit Test Framework

Most high-end computers adopt hybrid architecture, porting a large-scale scientific code onto accelerators is necessary. The paper presents a generic method for porting large-scale scientific code onto accelerators using compiler directives within a modularized function unit test platform. We have implemented the method and designed a software tool (SPEL) to port the E3SM Land Model (ELM) onto the GPUs in the Summit computer. SPEL automatically generates GPU-ready test modules for all ELM functions, such as CanopyFlux, SoilTemperature, and EcosystemDynamics. SPEL breaks the ELM into a collection of standalone unit test programs for easy code verification and further performance improvement. We further optimize several ELM test modules with advanced techniques, including memory reduction, reconstructed parallel loops, and asynchronous GPU kernel launch. We hope our study will inspire new toolkit developments that expedite large-scale scientific code porting with compiler directives.

Schwartz, Peter↗