Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Computational optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Analysis of Vector Particle-In-Cell (VPIC) memory usage optimizations on cutting-edge computer architectures

Vector Particle-In-Cell (VPIC) is one of the fastest plasma simulation codes in the world, with particle numbers ranging from one trillion on the first petascale system, Roadrunner, to ten trillion particles on the more recent Blue Waters supercomputer. As supercomputers continue to grow rapidly in size, so too does the gap between computing capability and memory capability. Current memory systems limit VPIC simulations greatly as the maximum number of particles that can be simulated directly depends on the available memory. In this study, we present a suite of VPIC memory optimizations (i.e., particle weight, half-precision, and fixed-point optimizations) that enable a significant increase in the number of particles in VPIC simulations. Here, we assess the optimizations’ impact on memory and runtime performance for a suite of cutting-edge computer architectures such has the NVIDIA V100 GPU, the IBM Power9, and the Fujitsu A64FX architectures. Our optimizations enable a 31.25% reduction in memory usage and up to 40% increase in the number of particles. This paper extends our work on developing particle storage format optimizations Tan et al.

97 MATHEMATICS AND COMPUTING↗

High-Fidelity and High-Performance Computational Simulations for Rapid Design Optimization of Sulfur Thermal Energy Storage (CRADA Report)

NREL and Element 16 collaborated on sulfur thermal energy storage modeling using NREL’s high performance computing (HPC) resources to assist its application in industrial processes. Industrial process heat (IPH) accounts for ~70% of US manufacturing energy use and is primarily produced by fossil fuel combustion. Approximately, 1500 TWht (~60% Terawatt hour thermal) of IPH demand is in the temperature range of 100-300°C. Industrial applications in this temperature range include drying, hydrothermal processing, thermal enhanced oil recovery, food and beverage, bioethanol production, etc. Cost-effective thermal energy storage (TES) that increases the utilization of waste and renewable heat (solar, geothermal, etc.) could provide significant energy savings and reliable heat sources, decrease emissions, and increase US manufacturing competitiveness through reductions in fuel consumption. This HPC4EI project facilitated Element 16’s development of low-cost and high-impact molten sulfur TES for dispatchable IPH. The development of a high-fidelity model validated by experimental data and HPC simulations enabled the successful resolution of the complex interplay between fluid dynamics and heat transfer processes during transient operation of sulfur TES, overcoming the numerical challenges posed by the non-linear temperature-dependent physical properties of sulfur. The project helped accelerate Element 16’s molten sulfur TES product design and support its broad applications.

25 ENERGY STORAGE↗

Enabling Scale-Up Through Multi-Fidelity Adaptive Computing

We present ideas from our ongoing work in adaptive computing - an optimization framework that allows us to strategically deploy various fidelity level experiments and simulations to guide decision making. The framework aims to enable uncertainty quantified scale-up of simulations and experiments, which causes increased complexity. A key feature of the framework is the integration of user-specified local model trustworthiness estimates. Adaptive sampling strategies allow us to optimally exploit the multiple fidelity level information and trustworthiness measures to arrive at the best decisions within a highly limited budget of objective function evaluations.

adaptive sampling↗

Multifunctional Chiral Chemically‐Powered Micropropellers for Cargo Transport and Manipulation

Practical applications of synthetic self-propelled nano and microparticles for microrobotics, targeted drug delivery, and manipulation at the nanoscale are rapidly expanding. However, fabrication limitations often hinder progress, resulting in relatively simple shapes and limited functionality. Here, taking advantage of 3D nanoscale printing, chiral micropropellers powered by the hydrogen peroxide reduction reaction are fabricated. Due to their chirality, the propellers exhibit multifunctional behavior controlled by an applied magnetic field: spinning in place (loitering), directed migration in the prescribed direction, capture, and transport of polymer cargo particles. Design parameters of the propellers are optimized by computation modeling based on mesoscale molecular dynamics. It is predicted by computer simulations, and confirmed experimentally, that clockwise rotating propellers attract each other and counterclockwise repel. These results shed light on how chirality and shape optimization enhance the functionality of synthetic autonomous micromachines.

36 MATERIALS SCIENCE↗

Optimal Control of the Energy-Saving Hybrid Hydraulic-Electric Architecture (HHEA) for Off-Highway Mobile Machines

Most off-highway constructions and agriculture equipment use hydraulics, which has unmatched power density, for power transmission and throttling as a means for control. A novel hybrid hydraulic-electric architecture (HHEA) has recently been proposed to improve efficiency for high-power machines that would have been cost-prohibitive to electrify directly. HHEA uses a set of common pressure rails (CPRs) to transmit the majority of power hydraulically and small electric motor drives to modulate that power and to achieve precise control. This article proposes a computationally efficient Lagrange multiplier method (LMM) for computing the optimal sequence of pressure rail selections to minimize energy use. This is needed to evaluate HHEA's energy-saving potential and for iterative architecture design and sizing. An interesting complication is that the cost function is not fully defined until the candidate control sequence is fully specified. This issue is dealt with by decomposing the original problem into a set of sub-problems with additional constraints that can be solved efficiently. Computational effort can be further reduced if actuators are optimized individually instead of together. However, additional steps are required to prevent the constraint functions from becoming discontinuous with respect to the Lagrange multipliers, which is necessary for meeting the constraints. Lastly, a case study of a construction machine demonstrates the efficacy of the method and shows that the HHEA reduces energy consumption by 68%-73% compared to the baseline load-sensing architecture.

Lagrange multiplier↗

Real-Time GPU-Accelerated OFDR With an Integrated Auxiliary Interferometer

A GPU-accelerated optical frequency domain reflectometry (OFDR) system with an improved integrated auxiliary interferometer is proposed. Unlike conventional approaches that require separate auxiliary interferometers and multiple detection channels, the proposed OFDR system embeds this functionality directly into the signal via an intentional beat component. This enables self-calibration of laser nonlinearity while maintaining a cost-effective hardware configuration. Building on this simplified configuration, the system leverages GPU acceleration with an NVIDIA RTX 4070 Ti to achieve real-time performance, delivering high-throughput signal processing for continuous OFDR interrogation. The signal processing pipeline comprises signal capture, resampling for nonlinearity compensation, and frequency shift computation, all optimized for parallel execution. Hardware benchmarking demonstrates substantial acceleration over CPU implementations, achieving up to a 45× speedup for resampling and frequency shift computations and enabling processing latencies below 30 ms. Thermal response validation is conducted under two complementary scenarios: localized heating using a water bath and cryogenic-temperature conditions using liquid nitrogen. Under localized heating, the system achieves an accuracy of 0.249 °C with a thermal sensitivity of 5.971 GHz/°C, while cryogenic-temperature validation demonstrates a frequency shift response with a sensitivity of 2.383 GHz/°C and an accuracy of 2.04 °C. The high acceleration of the proposed GPU-accelerated OFDR system and its accuracy are achieved by exploiting CUDA-based stride indexing, enabling efficient parallel segmentation and processing of large datasets without additional memory copies. The benchmarking results confirm the robustness, accuracy, and deployability of the proposed OFDR system across a wide temperature range, establishing it as a practical platform for real-time distributed fiber sensing in structurally dynamic environments.

Harb, Salah [Lawrence Berkeley National Laboratory↗

Personalized learning via task load optimization

A method for providing task load-optimized computer-generated training experiences to a user of a training system that includes: a display, a training simulator, a prediction program (ML1), and a training optimization program (ML2). In response to receiving a predicted optimal task load, ML2 provides a first training experience recommendation related to the training content and/or training conditions that, if utilized in providing a training experience to the user, is predicted to result in the predicted actual task load of the user equaling the predicted optimal task load. In response to receiving biometric information or performance metric information, ML1 determines the predicted actual task load. If the predicted actual task load does not match the predicted optimal task load, ML2 provides a second training experience recommendation and a second training experience is provided where at least one of the training content or the training conditions is changed.

Bertolli, Michael G.↗

Memory access optimization for particle operations in computational fluid dynamics-discrete element method simulations

Computational Fluid Dynamics - Discrete Element Method is used to model gas-solid systems in several applications in energy, pharmaceutical and petrochemical industries. Computational performance bottlenecks often limit the problem sizes that can be simulated at industrial scale. The data structures used to store several millions of particles in such large-scale simulations have a large memory footprint that does not fit into the processor cache hierarchies on current high-performance-computing platforms, leading to reduced computational performance. This paper specifically addresses this aspect of memory access bottlenecks in industrial scale simulations. The use of space-filling curves to improve memory access patterns is described and their impact on computational performance is quantified in both shared and distributed memory parallelization paradigms. The Morton space filling curve applied to uniform grids and k-dimensional tree partitions are used to reorder the particle data-structure thus improving spatial and temporal locality in memory. The performance impact of these techniques when applied to two benchmark problems, namely the homogeneous-cooling-system and a fluidized-bed, are presented. We report these optimization techniques lead to approximately two-fold performance improvement in particle focused operations such as neighbor-list creation and data-exchange, with ~ 1.5 times overall improvement in a fluidization simulation with 1.27 million particles.

97 MATHEMATICS AND COMPUTING↗

Multi-objective Bayesian optimization of ferroelectric materials with interfacial control for memory and energy storage applications

Optimization of materials’ performance for specific applications often requires balancing multiple aspects of materials’ functionality. Even for the cases where a generative physical model of material behavior is known and reliable, this often requires search over multidimensional function space to identify low-dimensional manifold corresponding to the required Pareto front. In this work, we introduce the multi-objective Bayesian optimization (MOBO) workflow for the ferroelectric/antiferroelectric performance optimization for memory and energy storage applications based on the numerical solution of the Ginzburg–Landau equation with electrochemical or semiconducting boundary conditions. MOBO is a low computational cost optimization tool for expensive multi-objective functions, where we update posterior surrogate Gaussian process models from prior evaluations and then select future evaluations from maximizing an acquisition function. Using the parameters for a prototype bulk antiferroelectric (PbZrO 3 ), we first develop a physics-driven decision tree of target functions from the loop structures. We further develop a physics-driven MOBO architecture to explore multidimensional parameter space and build Pareto-frontiers by maximizing two target functions jointly—energy storage and loss. This approach allows for rapid initial materials and device parameter selection for a given application and can be further expanded toward the active experiment setting. The associated notebooks provide both the tutorial on MOBO and allow us to reproduce the reported analyses and apply them to other systems (https://github.com/arpanbiswas52/MOBO_AFI_Supplements).

36 MATERIALS SCIENCE↗

Sequential Decision Making (SDM) for Mesh Refinement and Model Selection in Multiscale, Multi-Physics Applications

Intelligent automation and decision support are needed to enhance computational efficiency and robustness in multiscale and multi-physics problems, including materials science, manufacturing, and climate and weather modeling. Current scientific computing approaches for enabling decisions by scientists fail to explore the role of learning, reasoning, and probabilistic planning. Often these decisions are not performed in real-time during the computation but are made prior to the start of the computation, which must be interrupted in order to make changes to the prior choices. Such interruptions at different stages of the computation increase the total computing time and the need for a human expert to frequently monitor the results. State of art scientific computing methods consist of rule-based algorithms that cannot automatically adapt to a dynamically changing computing environment. The development of a Sequential Decision Making (SDM) framework will automate scientific computing by optimizing the policies for mesh refinement, time-stepping, model and algorithm selection, resource allocation, and pre and post-processing. Our agent SDM framework for scientific computing will consist of data-driven learning (Classifier), automated reasoning (contextual knowledge), and probabilistic planning (Reinforcement Learning). In this project, we focused on three problems to demonstrate our SDM framework on a set of ordinary and partial differential equations. Classification of Lorenz system regions using Feed-Forward Neural Networks examined learning in the SDM framework. On the other hand, reasoning and planning in the SDM framework were used in two problems: adaptive time-stepping for nonlinear ODEs using on-policy RL algorithms, and adaptive mesh refinement for 2-D PDEs using off-policy RL algorithms.

97 MATHEMATICS AND COMPUTING↗

Engagement: Hyperparameter Optimization of Generative Adversarial Network Models for High-Energy Physics Simulations

We present our SciDAC FASTMath-HEP partnership results for tuning generative adversarial models (GANs) for high energy physics applications. The GANs are used in hybrid simulations to accelerate otherwise time-consuming computations. We optimize for both, prediction accuracy and variability with the goal to find GAN architectures that are reliable and robust.

high energy physics↗

Computation of condition-dependent proteome allocation reveals variability in the macro and micro nutrient requirements for growth

Sustaining a robust metabolic network requires a balanced and fully functioning proteome. In addition to amino acids, many enzymes require cofactors (coenzymes and engrafted prosthetic groups) to function properly. Extensively validated resource allocation models, such as genome-scale models of metabolism and gene expression (ME-models), have the ability to compute an optimal proteome composition underlying a metabolic phenotype, including the provision of all required cofactors. Here we apply the ME-model for Escherichia coli K-12 MG1655 to computationally examine how environmental conditions change the proteome and its accompanying cofactor usage. We found that: (1) The cofactor requirements computed by the ME-model mostly agree with the standard biomass objective function used in models of metabolism alone (M-models); (2) ME-model computations reveal non-intuitive variability in cofactor use under different growth conditions; (3) An analysis of ME-model predicted protein use in aerobic and anaerobic conditions suggests an enrichment in the use of peroxyl scavenging acids in the proteins used to sustain aerobic growth; (4) The ME-model could describe how limitation in key protein components affect the metabolic state of E . coli . Genome-scale models have thus reached a level of sophistication where they reveal intricate properties of functional proteomes and how they support different E . coli lifestyles.

59 BASIC BIOLOGICAL SCIENCES↗

Optimized recursion relation for the computation of partition functions in the superconfiguration approach

Partition functions of a canonical ensemble of non-interacting bound electrons are a key ingredient of the super-transition-array approach to the computation of radiative opacity. A few years ago, we published a robust and stable recursion relation for the calculation of such partition functions. In this work, we propose an optimization of the latter method and explain how to implement it in practice. The formalism relies on the evaluation of elementary symmetric polynomials, which opens the way to further improvements.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Probing Accuracy-Speedup Tradeoff in Machine Learning Surrogates for Molecular Dynamics Simulations

The performance promise of machine learning surrogates of molecular dynamics simulations of soft materials is significant but generally comes at the cost of acquiring large training datasets to learn the complex relationships between input soft material attributes and output properties. Under the constraint of limited high-performance computing resources, optimizing the size of the training datasets becomes paramount. Using an artificial neural network based surrogate for molecular dynamics simulations of confined electrolytes, we explore the tradeoff between surrogate accuracy and computational gains. Accuracy is assessed by computing the root-mean-square errors between the surrogate predictions and the ground truth results obtained via molecular dynamics simulations. The computational performance is judged by evaluating the speedup which incorporates the training dataset creation time. Improvement in accuracy occurs with a loss of speedup, which scales as the inverse of the training dataset size. Furthermore, the link between surrogate generalizability and the accuracy-speedup tradeoff is assessed by examining the errors incurred in surrogate predictions on unseen, interpolated input variables and developing a net speedup metric to capture the associated gains.

Anions↗

Advocating Feedback Control for Human-Earth System Applications

This paper proposes a feedback control perspective for Human-Earth Systems (HESs) which essentially are complex systems that capture the interactions between humans and nature. Recent attention in HES research has been directed towards devising strategies for climate change mitigation and adaptation, aimed at achieving environmental and societal objectives. However, existing approaches heavily rely on HES models, which inherently suffer from inaccuracies due to the complexity of the system. Moreover, overly detailed models often prove impractical for optimization tasks. We propose a framework inheriting from feedback control strategies the robustness against model errors, because inaccuracies are mitigated using measurements retrieved from the field. The framework comprises two nested control loops. The outer loop computes the optimal inputs to the HES, which are then implemented by actuators controlled in the inner loop. Potential fields of applications are also identified and a numerical example is provided.

biological system modeling↗