Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Computational optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Development of Novel Ferritic-Martensitic Steels with Superior Creep Properties for Power Plant Applications

A creep resistant martensitic steel, CPJ-7, was developed with an operating temperature approaching 650°C. Subsequently, another creep resistant martensitic steel, JMP, was designed with the potential to operate at, or slightly above, 650°C. This report describes the development of these alloys from early iterations of CPJ steels to the final CPJ-7 formulation as well as the more heat resistant JMP steel formulation. The design originated from computational modeling for phase stability and precipitate strengthening using fifteen constituent elements. Approximately forty heats of CPJ and ten of JMP, each weighing ~7 kg, were vacuum induction melted. A computationally optimized heat treatment schedule was developed to homogenize the ingots prior to hot forging and rolling prior to final normalization and tempering. Overall, wrought and cast versions of CPJ-7 present superior creep properties when compared to wrought and cast versions of COST alloys for steam turbine and wrought and cast versions of P91/92 for boiler applications. For instance, the Larson Miller Parameter curve for CPJ-7 at 650°C almost coincides with that of COST E at 620°C. The prolonged creep life was attributed to slowing down the process of the destabilization of the MX and M 23 C 6 precipitates at 650°C. On the other hand, the cast version of CPJ-7 also revealed superior mechanical performance, well above commercially available cast 9% Cr martensitic steel or derivatives, especially for creep life at 650°C. The casting process employed slow cooling to simulate the conditions of a thick wall full-size steam turbine casing but utilized a separate homogenization step prior to final normalization and tempering. To advance the development of CPJ-7 for commercial applications, a process was used to scale up the production of the alloy using vacuum induction melting (VIM) and electroslag remelting (ESR), which underlined the importance of melt processing control of intentionally designed minor and trace elements in these advanced alloys. Following the work on CPJ-7, the JMP steels were designed with higher Co for increased solid solution strengthening, Si for oxidation resistance and increased W (with low Mo content) for matrix strength and stability as well as solid solution strengthening. The JMP steels showed increases in creep life compared to CPJ-7 between 118 to 150% at 650°C for testing at various stresses between 138 MPa and 207 MPa. On a Larson-Miller plot, the performance of the JMP steels surpasses that of state-of-the-art MARBN and other MARBN-type steels. The influence of various elements within the composition of the alloys on the microstructure and mechanical properties are discussed. This report presents approximately 420,000 h of in-house creep testing, the equivalent of almost 50 years of cumulative creep tests.

36 MATERIALS SCIENCE↗

Variation of γ′ Formers and Refractory Elements for Enhanced Creep resistance and Phase Stability of HAYNES® 282® Alloy

Modifications to the chemistry of alloy 282 were performed to improve the alloy’s resistance to long term creep deformation as well as phase stability. Alloys were prepared using vacuum induction melting, computationally optimized homogenization heat treatment and hot working. Phase stability studies were carried out for up to 5,000 hours at 800°C and 900°C while creep testing was performed at conditions leading to lives past 7,000 hours for temperatures ranging from 740°C to 900°C. The formation of σ and μ phases were reduced in the modified alloy. Atom probe tomography (APT) was performed on the specimens from the phase stability study to investigate changes in the elemental partitioning to the γ and γ′ phases. Post deformation microstructures were analyzed using EBSD and TEM to study the influence of the detrimental phases on damage accumulation during creep.

Detrois, Martin↗

Exact signed distance fields using parallel Fast Sweeping Method

Signed distance fields are often used in multiphysics simulations to track material interfaces. We present a simple methodology based on the fast sweeping method to generate the exact signed distance from triangular meshes and linear paths on Cartesian grids. The methodology propagates the closest primitive to the boundary to the rest of the domain following the characteristics. A local upwind criterion is used to decide between the new and existing closest primitive at each grid point while capturing the correct sign of the global function. The methodology has optimal computational complexity and runs efficiently in distributed-memory architectures. We include 2D and 3D test cases along with a resolution study up to 0.512 trillion zones and 1,000 computer cores. The solution strategy can also be applied to other types of meshes or collections of primitives.

97 MATHEMATICS AND COMPUTING↗

Multi-Scale Modeling and Prototype Development for Electrochemical CO2 Reduction (CRADA Final Report)

In this CRADA project, Lawrence Livermore National Laboratory, Stanford University, SLAC National Laboratory, and TotalEnergies collaboratively executed a multidisciplinary investigation of electrochemical reduction of CO2 to produce sustainable fuels and chemicals. Overall, the project led to an increased understanding of the fundamental processes involved in CO2 electrolysis, from the atomistic scale to the full electrolyzer device scale, ultimately leading to design guidelines for CO2 electrolyzers that will help in their future commercialization. As the model systems, Ag- and Cu-based catalysts were investigated in various forms depending on the electrochemical platform that was utilized to study the activity, selectivity, and durability towards electrochemical CO2 reduction. By employing experimental, theoretical, and computational techniques, the project team experimentally validated multi-physics models, evaluated the experimental levers that lead to increased electrolyzer reaction selectivity and energy efficiency, and used computational optimization to design higher performance electrodes. The learnings of this project were extensively documented in publicly available peer-reviewed journal publications and conference presentations, which serve as a foundation for further work to build from.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Choosing Transport Events for Initiating Splitting and Rouletting

A study was performed to determine which transport events should be used to initiate a weight window lookup to achieve the best variance reduction performance. A weight window lookup potentially triggers particle splitting (in important regions of phase space) or rouletting (in unimportant regions), thereby optimizing computational effort. Potential initiating transport events include collisions (both pre- and post-collision), geometry surface crossings, traversing a mean-free path, and streaming across a weight window boundary. Permutations of these initiating events were tested on an urban model with background radiation sources and a spent fuel cask with a neutron dose mesh tally. Generally, all methods perform better with finer weight window meshes. Tracking on weight windows performs well for coarse weight window meshes, while a combination of splitting each mean-free path, geometric surface crossing, and before collisions performs well for fine weight window meshes.

42 ENGINEERING↗

Optical design concept of the CMB-S4 large-aperture telescopes and cameras

CMB-S4 -- the next-generation ground-based cosmic microwave background (CMB) experiment - will significantly advance the sensitivity of CMB measurements and improve our understanding of the origin and evolution of the universe. CMB-S4 will deploy large-aperture telescopes fielding hundreds of thousands of detectors at millimeter wavelengths. We present the baseline optical design concept of the large-aperture CMB-S4 telescopes, which consists of two optical configurations: (i) a new off-axis, three-mirror, free-form anastigmatic design and (ii) the existing coma-corrected crossed-Dragone design. We also present an overview of the optical configuration of the array of silicon optics cameras that will populate the focal plane with 85 diffraction-limited optics tubes covering up to 9 degrees of field of view, up to $1.1 \, \rm mm$ in wavelength. We describe the computational optimization methods that were put in place to implement the families of designs described here and give a brief update on the current status of the design effort.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

CUDO: closed-form universal dwell-time optimization for computer-controlled optical surfacing

Precision optical figuring demands fast and accurate dwell time optimization to reach nanometer- and sub-nanometer-level accuracy in next-generation optical systems. We introduce CUDO (closed-form universal dwell-time optimization), the first, to the best of our knowledge, unified closed-form analytical framework that supports both function-form and matrix-form dwell time models in computer-controlled optical surfacing (CCOS). In contrast to traditional methods, which rely on iterative optimization and hyperparameter tuning, our framework derives direct analytical solutions with no adjustable parameters. This approach unifies the solution principles of existing methods within a single mathematical model, delivering three key advantages: (1) accuracy on par with, or superior to, iterative solvers, (2) substantial reduction in computation time, and (3) numerical robustness. Comparative studies with prior art confirm that closed-form solutions achieve equivalent residual error while removing runtime bottlenecks. By simplifying the implementation and enabling real-time, scalable deployment, CUDO establishes a practical foundation for future deterministic fabrication of large-aperture and high-performance optics.

36 MATERIALS SCIENCE↗

Optimization of distributed compute resources utilization in the CMS Global Pool

The CMS Submission Infrastructure is the primary system for managing computing resources for CMS workflows, including data processing, simulation, and analysis. It integrates geographically distributed resources from Grid, HPC, and cloud providers into federated pools managed by HTCondor and Glidein- WMS, for a total of around 500k CPU cores. This system dynamically manages workloads based on priorities defined by the collaboration. Additionally, CMS scheduling strategies must be flexible to handle multiple concurrent workloads while considering changing processing demands and resource availability from various providers.Efficient utilization of vast amounts of distributed compute resources is a key element for the success of the scientific programs of the LHC experiments. Optimizing the system is essential to maximize resource efficiency and fully utilize the distributed computing power. The CMS Submission Infrastructure team thus systematically investigates sources of inefficiency in workload scheduling to reduce their impact. In addition, a strategy of pilot overloading has been introduced to compensate for other inefficiency sources, thereby optimizing resource utilization and enhancing computational throughput.

Mascheroni, Marco [UC, San Diego (main)]↗

Finding inputs that trigger floating-point exceptions in heterogeneous computing via Bayesian optimization

Testing code for floating-point exceptions is crucial as exceptions can quickly propagate and produce unreliable numerical answers. The state-of-the-art to test for floating-point exceptions in heterogeneous systems is quite limited and solutions require the application’s source code, which precludes their use in accelerated libraries where the source is not publicly available. We present an approach to find inputs that trigger floating-point exceptions in black-box CPU or GPU functions, i.e., functions where the source code and information about input bounds are unavailable. Our approach is the first to use Bayesian optimization (BO) to identify such inputs and uses novel strategies to overcome the challenges that arise in applying BO to this problem. Here, we implement our approach in the XSCOPE framework and demonstrate it on 58 functions from the CUDA Math Library and 81 functions from the Intel Math Library. XSCOPE is able to identify inputs that trigger exceptions in about 73% of the tested functions.

97 MATHEMATICS AND COMPUTING↗

Communication Lower Bounds and Optimal Algorithms for Multiple Tensor-Times-Matrix Computation

Multiple tensor-times-matrix (Multi-TTM) is a key computation in algorithms for computing and operating with the Tucker tensor decomposition, which is frequently used in multidimensional data analysis. Here, we establish communication lower bounds that determine how much data movement is required (under mild conditions) to perform the Multi-TTM computation in parallel. The crux of the proof relies on analytically solving a constrained, nonlinear optimization problem. We also present a parallel algorithm to perform this computation that organizes the processors into a logical grid with twice as many modes as the input tensor. We show that, with correct choices of grid dimensions, the communication cost of the algorithm attains the lower bounds and is therefore communication optimal. Finally, we show that our algorithm can significantly reduce communication compared to the straightforward approach of expressing the computation as a sequence of tensor-times-matrix operations when the input and output tensors vary greatly in size.

HBL-inequalities↗

Computational Analysis and Optimized Modeling of Geomagnetically Induced Currents in Power Transformers

In this project we aim to better understand the effect of geomagnetically induced currents (GIC) on power transformers. Expanding upon our previous work focused on producing a methodology for accuracy-enhanced computation of GIC signatures (i.e., time-domain current magnitude variation for the event duration) from a combination of physics-based and data-driven computational tools, we propose the use of these GIC signatures as inputs for a physically-detailed and optimized model of the power transformer to investigate how GIC determination and transformer modeling influence the evaluation of GIC effects on the transformer operation, as well as in its interaction with the power grid.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Enhancing the Functionality of a Hollow Scaffold Solid State Bioreactor via Computer Aided Design Optimization

The concern over greenhouse gases, methane (CH 4 ) and carbon dioxide (CO 2 ), is increasing rapidly. There have been strides to find solutions to this global issue but there is not a clear path to a successful end goal. The concentration of CH 4 and CO 2 in the atmosphere has increased significantly over the last 60 years, methane is a great source of concern due to its ability to trap a high amount of heat in the atmosphere. These greenhouse gases contribute to global warming which has caused changes in the environment, including the melting of ice caps, and altered weather patterns. Solutions for these pressing challenges have led to different avenues of methane mitigation one of which is the development of solid-state bioreactors. These reactors harness the power of biological species that have evolved to use methane as an energy source. The development and optimization of Hollow Scaffold Solid State Bioreactors (HS-SSBR) has become readily available due to the advancements in additive manufacturing technology and accessibility of computer aided design (CAD) software. With laboratory scale experiments, time and effort are of great importance. Enhancing the design of the HS-SSBR to create a more user-friendly interface, but also increase the functionality of the reactor. The reactor's design improvements focus on better dispersion of methane and circulation of media.

36 MATERIALS SCIENCE↗

Development and Analysis of Optimal Multilevel Solvers on Advanced Computers. Final Report

Constrained optimization in the context of time dependent, partial differential equations (PDE) leads to a symmetric, block-tridiagonal system of nonlinear equations that must be solved repeatedly in an iterative solution strategy. The blocks represent spatial discretization, while the connection between the blocks represents a forward and backward integration in time. The focus of this project is to apply a parallel-in-time (PiT) solution technique to the large block-triangular system.

97 MATHEMATICS AND COMPUTING↗

SymProp: Scaling Sparse Symmetric Tucker Decomposition via Symmetry Propagation

Sparse symmetric tensors are an important class of tensors, and their decompositions serve as powerful tools for revealing low-rank structures. This paper introduces SymProp, a novel approach for scaling sparse symmetric Tucker decomposition by propagating symmetry through intermediate computations. SymProp optimizes two key computational kernels: Sparse Symmetric Tensor Times Same Matrix chain (S3 TTMc) for Higher-Order Orthogonal Iteration (HOOI) and Sparse Symmetric Tensor Times Same Matrix chain Times Core (S3 TTMcTC) for Higher-Order QR Iteration (HOQRI). Our method employs a metaprogramming-based index iteration approach to efficiently handle the upper triangular parts of intermediate dense symmetric tensors. SymProp achieves up to 50.9× speedup over SPLATT and up to 360.8× over Compressed Sparse Symmetric (CSS) format on the S3 TTMc operation. Moreover, our S3 TTMc and S3 TTMcTC implementations support tensor orders four levels higher than state-of-the-art methods. Our HOQRI demonstrates superior scalability and up to a 33.6× speedup over optimized HOOI. By enabling more scalable Tucker decompositions for higher orders, decomposition ranks, and dimension sizes, SymProp opens new possibilities for analyzing complex hypergraph structures in fields such as network science, data mining, and machine learning.

Li, Zecheng [North Carolina State University]↗

Muon track reconstruction in a segmented bolometric array using multi-objective optimization

Recent advances in segmented solid-state detector arrays for rare-event searches have allowed the technology to approach the ton-scale in detector mass and the scale of meters in size. Often focused around searches for neutrinoless double-beta decay or direct dark matter detection, such experiments also have the capability to search for exotic particles that leave track-like signatures across their volume. However, the segmented nature of such detector arrays often sets the spatial resolution and makes the problem of reconstructing track-like paths non-trivial. Here, in this paper, we present an algorithm that improves reconstruction of track-like events in segmented detectors using multi-objective optimization — a computational technique that optimizes more than one cost function at a time without specifying a quantitative weighting between them. Such a technique allows the reconstruction of tracks through a detector and the determination of path-lengths through individual elements. When combined with the reconstructed energy depositions in each element this allows for a calculation of the stopping power of track-like particles and opens the door to searches for particles with abnormal stopping power like monopoles or lightly-ionizing particles (LIPs). Results are presented which evaluate the precision of the reconstruction tools as they currently stand against Monte Carlo generated data. The algorithm is presented in the context of the CUORE experiment, but has applications to other segmented calorimeter detectors.

47 OTHER INSTRUMENTATION↗