Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “batch size”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Universal method for the optimization of HDC coating uniformity on non-planar, non-stationary substrates for inertial confinement fusion targets

The thickness uniformity of chemical vapor deposited (CVD) diamond coatings on non-planar, non-stationary substrates depends on both the intrinsic instantaneous coating thickness distribution (ICTD) of the coating conditions used and, if applicable, on the frequency of substrate reorientation. While important for many CVD diamond applications, the relative impact of the ICTD and substrate reorientation on the coating thickness uniformity has not been studied. In this work, we systematically investigate the effect of these factors for microwave-plasma chemical vapor deposition (MPCVD) of diamond (referred to as high density carbon (HDC) in the inertial confinement fusion (ICF) community) coatings on spherical, rolling substrates. This coating technique is used to fabricate capsules for ICF experiments, which require extreme coating uniformity with <0.3 % thickness variation (so-called Mode 1 or M1) to ensure symmetric compression of imploding targets. To extract the otherwise unobservable reorientation timescale (Δt), Monte Carlo simulations were performed using experimental ICTD data as input. This combined approach confirms scaling relationships between the substrate reorientation timescale as well as coating thickness and coating uniformity, as expected from a 3D random walk. Simulations confirm that M1 is Rayleigh-distributed and scales as (Δt) 1/2 , consistent with the randomization of two angles that determine orientation of a sphere. We also demonstrate that, under the conditions studied, Δt is the dominant factor in determining thickness uniformity while the intrinsic ICTD has minimal impact. Finally, experiments show that Δt can be affected by total batch size under constant agitation conditions due to space constraints that limit the capsule reorientation kinetics. In conclusion, this study highlights the utility of a combined experiment-simulation approach as a general methodology for understanding and improving coating uniformity on non-planar, non-stationary substrates.

Capsule↗

A strategy for automated core design to increase economic viability and minimize fuel fragmentation, relocation, and dispersal susceptibility in high-burnup cores

The nuclear industry aims to increase the cycle length of pressurized water reactors from 18 to 24 months to increase power plant capacity factors and economic viability. These cycle length extensions will inherently require fuel rods to exceed the current peak rod average burnup limit of 62 GWd/MTU. A chief concern of operating beyond the current burnup limit is the fuel fragmentation, relocation, and dispersal (FFRD) phenomenon in which pulverized fuel fragments can axially relocate and escape through a burst in the cladding formed during a loss-of-coolant accident. In this work, we demonstrate an approach for automating core design employing an optimization tool based on a penalty-free, parallel simulated annealing algorithm to produce pressurized water reactor core designs with two different optimization objectives. The two objectives were to produce core designs with (1) mitigated FFRD susceptibility while achieving 24-month cycle lengths (2) maximum cycle length with no regard for the likelihood of FFRD. Batch size was considered in tandem with both cases to maximize economic viability. The PARCS nodal model was the primary reactor physics tool used in the optimizations and used nuclear cross sections calculated with 2D Polaris lattice physics models. Reactor performance and safety characteristics of the optimized cores were verified using high-fidelity Virtual Environment for Reactor Applications models. The core designs produced by the optimization tool are compared with each other and to a high-burnup core design produced and analyzed in previous works to highlight the fuel management strategies that may enhance high-burnup reactor safety and economic viability. The optimized cores satisfied their respective objective functions, producing a maximum cycle length of 720 effective full-power days in one core design and one that may reduce FFRD susceptibility by up to 50% based on the first-order approximation to FFRD risk formulated in this work. The optimized cores met most constraints but exceeded the hot channel factor limit, especially in FFRD cases where fresh fuel carried more power. Furthermore, this highlights the need for future lattice-level optimizations and broader assembly options.

Cycle length↗

Implementation of McMurchie–Davidson Algorithm for Gaussian AO Integrals Suited for SIMD Processors

We report an implementation of the McMurchie− Davidson evaluation scheme for 1- and 2-particle Gaussian AO integrals designed for processors with Single Instruction Multiple Data (SIMD) instruction sets. Like in our recent MD implementation for graphical processing units (GPUs) [Asadchev, A.; Valeev, E. F.. J. Chem. Phys. 2024, 160, 244109.], variable-sized batches of shellsets of integrals are evaluated at a time. By optimizing for the floating point instruction throughput rather than minimizing the number of operations, this approach achieves up to 50% of the theoretical hardware peak FP64 performance for many common SIMD-equipped platforms (AVX2, AVX512, NEON), which translates to speedups of up to 30 over the state-of-the-art one-shellset-at-a-time implementation of Obara−Saika-type schemes in Libint for a variety of primitive and contracted integrals. As with our previous work, we rely on the standard C++ programming language such as the std::simd standard library feature to be included in the 2026 ISO C++ standard without any explicit code generation to keep the code base small and portable. The implementation is part of the open source LibintX library freely available at https://github.com/ValeevGroup/libintx.

Basis sets↗

Glass-Bonded Monazite Waste Forms for Lanthanide and Actinide Immobilization: From Theoretical Design to Scale-Up Production and Characterization

The development of nuclear waste forms for both existing and future nuclear wastes is critical to ensuring global environmental safety. This study focuses on waste management from molten salt reactors, where fuel exists in a salt form and could be processed in real time for the removal of neutron poisons such as xenon isotopes (e.g., 135 Xe) and rare earth elements (REEs, e.g., 149 Sm). To ensure safe, stable, and long-term disposal in geological repositories, REEs must be incorporated into a durable waste form. Iron-phosphate glasses are a promising candidate due to their low melting points, high chemical durability, and their ability to incorporate high concentrations of REEs. In this study, we successfully prepared iron-phosphate glass waste forms with high Nd loadings (up to 37 mass %) in batch sizes ranging from small (23 g) to large (1600 g). The resulting materials contained up to 75 mass % NdPO 4 , contributing to their mechanical resilience and exceptional chemical durability. These findings highlight the potential of iron-phosphate glasses as high-efficiency, chemically durable waste forms and demonstrate the successful transition from theoretical design to scaled-up production.

amorphous materials↗

Scalable nanomanufacturing of chalcogenide inks: a case study on thermoelectric V–VI nanoplates

Solution-processed semiconducting main-group chalcogenides (MMCs) have attracted increasing research interest for next-generation device technologies owing to their unique nanostructures and superior properties. To achieve the full potential of MMCs, the development of highly universal, scalable, and sustainable synthesis and processing methods of chalcogenide particles is thus becoming progressively more important. Here we studied scalable factors for the synthesis of two-dimensional (2D) V–VI chalcogenide nanoplates (M 2 Q 3 : M = Sb, Bi; Q = Se, Te) and systematically investigated their colloidal behaviour and chemical stability. Based on a solvent engineering technique, we demonstrated scale-up syntheses of MMCs up to a 900% increase of batch size compared with conventional hydrazine-based gram-level syntheses, and such a scalable approach is highly applicable to various binary and ternary MMCs. Furthermore, we studied the stability of printable chalcogenide nanoparticle inks with several formulation factors including solvents, additives, and pH values, resulting in inks with high chemical stability (>4 months). As a proof of concept, we applied our solution-processed chalcogenide particles to multiple additive manufacturing methods, confirming the high printability and processability of MMC inks. Furthermore, the ability to combine the top-down designing freedom of additive manufacturing with bottom-up scalable synthesis of chalcogenide particles promises great opportunities for large-scale design and manufacturing of chalcogenide-based functional devices for broad application.

36 MATERIALS SCIENCE↗

Real-time semantic segmentation on FPGAs for autonomous vehicles with hls4ml

In this paper, we investigate how field programmable gate arrays can serve as hardware accelerators for real-time semantic segmentation tasks relevant for autonomous driving. Considering compressed versions of the ENet convolutional neural network architecture, we demonstrate a fully-on-chip deployment with a latency of 4.9 ms per image, using less than 30% of the available resources on a Xilinx ZCU102 evaluation board. The latency is reduced to 3 ms per image when increasing the batch size to ten, corresponding to the use case where the autonomous vehicle receives inputs from multiple cameras simultaneously. We show, through aggressive filter reduction and heterogeneous quantization-aware training, and an optimized implementation of convolutional layers, that the power consumption and resource utilization can be significantly reduced while maintaining accuracy on the Cityscapes dataset.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Reducing Communication Overhead in Federated Learning for Network Anomaly Detection with Adaptive Client Selection

Communication overhead in federated learning (FL) poses a significant challenge for network anomaly detection systems, where the myriad of client configurations and network conditions can severely impact system efficiency and detection accuracy. While existing approaches attempt to address this through individual optimization techniques, they often fail to maintain the delicate balance between reduced overhead and detection performance. This paper presents an adaptive FL framework that dynamically combines batch size optimization, client selection, and asynchronous updates to achieve efficient anomaly detection. Through extensive profiling and experimental analysis on two distinct datasets-UNSW-NBIS for general network traffic and ROAD for automotive networks-our framework reduces communication overhead by 97.6%; (from 700.0s to 16.8s) compared to synchronous baseline approaches while maintaining comparable detection accuracy (95.10%; vs. 95.12%;). Statistical validation using Mann-Whitney U test confirms significant improvements (p < 0.05) over existing FL approaches across both datasets, demonstrating the framework's adaptability to different network security contexts. Detailed profiling analysis reveals the efficiency gains through dramatic reductions in GPU operations and memory transfers while maintaining robust detection performance under varying client conditions.

Marfo, William [University of Texas at El Paso]↗

MDLoader: A Hybrid Model-Driven Data Loader for Distributed Graph Neural Network Training

Scalable data management is essential for processing large scientific dataset on HPC platforms for distributed deep learning. In-memory distributed storage is preferred for its speed, enabling rapid, random, and frequent data access required by stochastic optimizers. Processes use one-sided or collective communication to fetch remote data, with optimal performance depending on (i) dataset characteristics, (ii) training scale, and (iii) interconnection network. Empirical analysis shows collective communication excels with larger mini-batch sizes and/or fewer processes, whereas one-sided communication outperforms at larger scales. We propose MDLoader, a hybrid in-memory data loader for distributed graph neural network training. MDLoader features a model-driven performance estimator that dynamically selects between one-sided and collective communication at the beginning of training using Tree of Parzen Estimators (TPE). Evaluations on NERSC Perlmutter and OLCF Summit show MDLoader outperforms single-backend loaders by up to 2.83 × and predicts the suitable communication method with 96.3% (Perlmutter) and 94.3% (Summit) success rate.

Bae, Jonghyun↗

Fabrication of Six Degrees-of-Freedom Hexflex Positioner With Integrated Strain Sensing Using Nonlithographically Based Microfabrication

In this study, a process flow is described for the low cost, flexible fabrication of metal micro-electromechanical systems (MEMS) with high performance integrated sensing. The process is capable of producing new designs in ≈1 week at an average unit cost of <$1 k/device even at batch sizes of ≈1–10, with expected sensing performance limits of about 135 dB over a 10 kHz sensor bandwidth. This is a ≈20× reduction in cost, ≈25× reduction in time, and potentially >30× increase in sensing dynamic range over comparable state-of-the-art compliant nanopositioners. The nonlithographically based microfabrication (NLBM) process is uniquely suited to create high performance nanopositioning architectures which are customizable to the positioning requirements of a range of nanoscale applications. These can significantly reduce the cost of nanomanufacturing research and development, as well as accelerate the development of new processes and the testing of fabrication process chains without excess capital investment. A six degrees-of-freedom (6DOF) flexural nanopositioner with integrated sensing for all 6DOF was fabricated using the newly developed process chain. The fabrication process was measured to have ≈30 μ m alignment. Sensor arm, flexure, and trace widths of 150 μ m, 150 μ m, and 800 μ m, respectively, were demonstrated. Process capabilities suggest lower bounds of 25 μ m, 50 μ m, and 100 μ m, respectively. Dynamic range sensing of 52 dB was demonstrated for the nanopositioner over a 10 kHz sensor bandwidth. Improvements are proposed to approach sensor performance of about 135 dB over a 10 kHz sensor bandwidth.

42 ENGINEERING↗

Efficient Training of Deep Neural Operator Networks via Randomized Sampling

Neural operators (NOs) employ deep neural networks to learn the mappings between infinitedimensional function spaces. Deep operator network (DeepONet), a popular NO architecture, has demonstrated success in the real-time prediction of complex dynamics across various scientific and engineering applications. In this work, we introduce a random sampling technique to be adopted during the training of DeepONet, aimed at improving the generalization ability of the model, while significantly reducing the computational time. The proposed approach targets the trunk network of the DeepONet model that outputs the basis functions corresponding to the spatiotemporal locations of the bounded domain on which the physical system is defined. While constructing the loss function, DeepONet training traditionally considers a uniform grid of spatiotemporal points at which all the output functions are evaluated for each iteration. This approach leads to a larger batch size, resulting in poor generalization and increased memory demands, due to the limitations of the stochastic gradient descent (SGD) optimizer. The proposed random sampling over the inputs of the trunk net mitigates these challenges, improving generalization and reducing the memory requirements during training, resulting in significant computational gains. We validate our hypothesis through three benchmark examples, demonstrating substantial reductions in training time while achieving comparable or lower overall test errors relative to the traditional training approach. Our results indicate that incorporating randomization in the trunk network inputs during training enhances the efficiency and robustness of DeepONet, offering a promising avenue for improving the framework’s performance in modeling complex physical systems.

Karumuri, Sharmila [Department of Civil & Systems ↗

Data-Driven Protection Software to classify fault locations by protective zone in distribution systems with high PV penetration

The software contains (a) the source codes to generate Point-on-Wave (PoW) transient data for any feeder model in Alternative Transient Program (ATP) format. Codes provide options to change different steady state settings, including the loading condition and PV capacity and transient state setting like faults type, location and initiation time (b) data post-processing source code to converted data from native format to COMTRADE, csv, HDF5 (c) Docker container to train CNN to classify fault locations by protective zone. The container takes dataset and other training parameters (sampling rate, training epochs, batch size etc) as input to train CNN. The container writes back the trained CNN model, training and testing metrics and plots to the local workstation

Ramesh, Meghana↗

A graphics processing unit accelerated sparse direct solver and preconditioner with block low rank compression

We present the GPU implementation efforts and challenges of the sparse solver package STRUMPACK. The code is made publicly available on github with a permissive BSD license. STRUMPACK implements an approximate multifrontal solver, a sparse LU factorization which makes use of compression methods to accelerate time to solution and reduce memory usage. Multiple compression schemes based on rank-structured and hierarchical matrix approximations are supported, including hierarchically semi-separable, hierarchically off-diagonal butterfly, and block low rank. Here, in this paper, we present the GPU implementation of the block low rank (BLR) compression method within a multifrontal solver. Our GPU implementation relies on highly optimized vendor libraries such as cuBLAS and cuSOLVER for NVIDIA GPUs, rocBLAS and rocSOLVER for AMD GPUs and the Intel oneAPI Math Kernel Library (oneMKL) for Intel GPUs. Additionally, we rely on external open source libraries such as SLATE (Software for Linear Algebra Targeting Exascale), MAGMA (Matrix Algebra on GPU and Multi-core Architectures), and KBLAS (KAUST BLAS). SLATE is used as a GPU-capable ScaLAPACK replacement. From MAGMA we use variable sized batched dense linear algebra operations such as GEMM, TRSM and LU with partial pivoting. KBLAS provides efficient (batched) low rank matrix compression for NVIDIA GPUs using an adaptive randomized sampling scheme. The resulting sparse solver and preconditioner runs on NVIDIA, AMD and Intel GPUs. Interfaces are available from PETSc, Trilinos and MFEM, or the solver can be used directly in user code. We report results for a range of benchmark applications, using the Perlmutter system from NERSC, Frontier from ORNL, and Aurora from ALCF. For a high frequency wave equation on a regular mesh, using 32 Perlmutter compute nodes, the factorization phase of the exact GPU solver is about 6.5× faster compared to the CPU-only solver. The BLR-enabled GPU solver is about 13.8× faster than the CPU exact solver. For a collection of SuiteSparse matrices, the STRUMPACK exact factorization on a single GPU is on average 1.9× faster than NVIDIA’s cuDSS solver.

97 MATHEMATICS AND COMPUTING↗

Shielded and Low-Damping SPM Probes For Quantitative Electrical and Electrochemical Characterization (Phase IIB)

The purpose of this DOE SBIR award was to develop new scanning probe microscopy tools for the elucidation of electrical and electrochemical properties at the nanoscale. High performance probes are important for energy materials research. In both solar cells and batteries, the essential physics and chemistry occur at the nanoscale and nanoscale characterization tools are required for their elucidation. The introduced new models are a shielded probe and an electrically insulated probe where only less than one micron of a conductive tip is exposed. In Phases I and IIa, we explored the development of cost-effective methods to insulate cantilevers from aqueous solutions and patented a technology that provides an insulated electrical contact to the tip. Also, our technology was adapted to be compatible with all popular commercial AFM systems. In Phase IIb, we improved reproducibility and batch size to drive down costs while maintaining quality control with testing. We transformed our product into a convenient turnkey solution to make electrochemical characterization approachable to a wide range of researchers. In January 2021, the technology was well received by the AFM community and Nanosurf Inc. (Woburn, MA) acquired Scuba Probe Technologies. The technology and expertise of Scuba Probe Technologies complements Nanosurf’s development of the full range of nano-electrical characterization tools.

36 MATERIALS SCIENCE↗

Enabling real-time adaptation of machine learning models at x-ray Free Electron Laser facilities with high-speed training optimized computational hardware

The emergence of novel computational hardware is enabling a new paradigm for rapid machine learning model training. For the Department of Energy’s major research facilities, this developing technology will enable a highly adaptive approach to experimental sciences. In this manuscript we present the per-epoch and end-to-end training times for an example of a streaming diagnostic that is planned for the upcoming high-repetition rate x-ray Free Electron Laser, the Linac Coherent Light Source-II. We explore the parameter space of batch size and data parallel training across multiple Graphics Processing Units and Reconfigurable Dataflow Units. We show the landscape of training times with a goal of full model retraining in under 15 min. Although a full from scratch retraining of a model may not be required in all cases, we nevertheless present an example of the application of emerging computational hardware for adapting machine learning models to changing environments in real-time, during streaming data acquisition, at the rates expected for the data fire hoses of accelerator-based user facilities.

97 MATHEMATICS AND COMPUTING↗

Adaptive, Active Learning, and Multifidelity Monte Carlo Methods in the MOOSE Stochastic Tools Module

MOOSE is an open-source computational platform for constructing multi-physics models and executing them in a massively parallel fashion. It has a stochastic tools module (STM) for forward/inverse uncertainty quantification (UQ) and surrogate modeling. This presentation details some recent developments to the STM with respect to the implementation of adaptive, active learning, and multifidelity Monte Carlo methods for forward UQ of computational models. Specifically, the adaptive Monte Carlo methods include Markov Chain Monte Carlo (MCMC)-driven algorithms like adaptive importance sampling and parallelized subset simulation for statistical QoI estimation, rare events analysis, and stochastic gradient-free optimization. The active learning methods include Gaussian Process (GP) surrogates and their training via Adam optimization, design of acquisition functions, and integration with samplers like Monte Carlo, adaptive importance, and parallelized subset simulation. These active learning methods are also designed to work in a batch mode, wherein, the required calls to the full computational model are executed in parallel whenever a user-specified batch size is met. The multifidelity methods in STM are broadly divided into two categories: hierarchical, where a defined hierarchy exists among the low-fidelity models, and peer, where all the low-fidelity models are treated equally. A GP surrogate is used to learn the differences between the low- and high-fidelity models in both multifidelity categories, and acquisition functions from the active learning classes are used to decide whether to rely on a low-fidelity model or call the expensive high-fidelity model. Alongside the software description and usage, applications are also presented to nuclear engineering computational models including a TRISO nuclear fuel particle, a reactor pressure vessel, and a heat-pipe microreactor.

97 MATHEMATICS AND COMPUTING↗

MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models

Mixture of Experts (MoE) models have enabled the scaling of Large Language Models (LLMs) and Vision Language Models (VLMs) by achieving massive parameter counts while maintaining computational efficiency. However, MoEs introduce several inference-time challenges, including load imbalance across experts and the additional routing computational overhead. To address these challenges and fully harness the benefits of MoE, a systematic evaluation of hardware acceleration techniques is essential. We present MoE-Inference-Bench, a comprehensive study to evaluate MoE performance across diverse scenarios. We analyze the impact of batch size, sequence length, and critical MoE hyperparameters such as FFN dimensions and number of experts on throughput. We evaluate several optimization techniques on Nvidia H100 GPUs, including pruning, Fused MoE operations, speculative decoding, quantization, and various parallelization strategies. Our evaluation includes MoEs from the Mixtral, DeepSeek, OLMoE and Qwen families. The results reveal performance differences across configurations and provide insights for the efficient deployment of MoEs.

Chitty-Venkata, Krishna Teja↗

Component-Level Inverse Design of Transmon Qubits Using Neural Networks

Designing a superconducting qubit to realize specific Hamiltonian parameters typically requires iterating through a time and compute-intensive forward loop in which the designer chooses a layout geometry, simulates it, extracts circuit parameters such as capacitances, and refines the geometry. We study the inverse version of this task using a neural-network workflow that maps target Hamiltonian parameters directly to component-level layout parameters, which we subsequently demonstrate on a planar transmon layout. During training, we pair the inverse model with a frozen forward surrogate model and evaluate the loss in Hamiltonian space rather than in layout-parameter space. In validation against a conventional EM solver, 97% of generated designs produce usable geometries, and the inverse-plus-surrogate pipeline reaches mean percent errors of 0.73% for qubit frequency and 1.58% for anharmonicity, comparable to or below the fabrication and simulation-to-measurement uncertainty expected for academic-process transmon devices of this type. A single pipeline query takes ~60 ms on CPU, versus ~2 min for a conventional EM capacitance extraction on the same hardware, a speedup of approximately 2,000x. Batching minimizes the AI model inference overhead, reducing the runtime to 3.1 microseconds per sample on CPU and 2.6 microseconds per sample on GPU at a batch size of 2048, resulting in speedups of 3.9 x 10^7 and 4.6 x 10^7, respectively, relative to a single conventional CPU EM extraction. Our results indicate that component-level inverse design usefully extends and complements conventional EM simulation, including for small datasets on the order of 1,000 samples.

Seidel, Olivia [Fermilab; Texas U., Arlington]↗