Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “batch size”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Sensitivity Study of Mini-Batch Size on a Long Short-Term Memory Network for In-situ Sensing of Core-to-shell Ratio of Microencapsulated Phase Change Materials

Microencapsulated phase change materials are being studied for applications for thermal energy storage in concentrated solar fields. During fabrication, the thickness of the encapsulation cannot be readily measured for real-time control. Therefore, a machine learning network, specifically a Long Short-Term Memory network, is being developed to estimate the ratio of the shell radius to core radius based on a one second temperature history. The mini-batch size determines how often the algorithm weights are updated during network training, and shuffle indicates whether the training data is shuffled during training. A general factorial design is used to analyze the effects of varying mini-batch size and shuffle, along with the core-to-shell ratio, on the RMSE of the response from the Long Short-Term Memory network. It was found that the network performed better for smaller core to shell ratios (less than 0.6) and had the lowest RMSE when the minibatch size was 128. The minimum RMSE found was 0.00501.

Shannon, Rebecca↗

SineKAN: Kolmogorov-Arnold Networks using sinusoidal activation functions

Recent work has established an alternative to traditional multi-layer perceptron neural networks in the form of Kolmogorov-Arnold Networks (KAN). The general KAN framework uses learnable activation functions on the edges of the computational graph followed by summation on nodes. The learnable edge activation functions in the original implementation are basis spline functions (B-Spline). Here, we present a model in which learnable grids of B-Spline activation functions are replaced by grids of re-weighted sine functions (SineKAN). We evaluate numerical performance of our model on a benchmark vision task. We show that our model can perform better than or comparable to B-Spline KAN models and an alternative KAN implementation based on periodic cosine and sine functions representing a Fourier Series. Further, we show that SineKAN has numerical accuracy that could scale comparably to dense neural networks (DNNs). Compared to the two baseline KAN models, SineKAN achieves a substantial speed increase at all hidden layer sizes, batch sizes, and depths. Current advantage of DNNs due to hardware and software optimizations are discussed along with theoretical scaling. Additionally, properties of SineKAN compared to other KAN implementations and current limitations are also discussed.

Reinhardt, Eric↗

Scale-up Unlearnable Examples Learning with High-performance Computing

Recent advancements in AI models, like ChatGPT, are structured to retain user interactions, which could inadvertently include sensitive healthcare data. In the healthcare field, particularly when radiologists use AI-driven diagnostic tools hosted on online platforms, there is a risk that medical imaging data may be repurposed for future AI training without explicit consent, spotlighting critical privacy and intellectual property concerns around healthcare data usage. Addressing these privacy challenges, a novel approach known as Unlearnable Examples (UEs) has been introduced, aiming to make data unlearnable to deep learning models. A prominent method within this area, called Unlearnable Clustering (UC), has shown improved UE performance with larger batch sizes but was previously limited by computational resources (e.g., a single workstation). To push the boundaries of UE performance with theoretically unlimited resources, we scaled up UC learning across various datasets using Distributed Data Parallel (DDP) training on the Summit supercomputer. Our goal was to examine UE efficacy at high-performance computing (HPC) levels to prevent unauthorized learning and enhance data security, particularly exploring the impact of batch size on UE’s unlearnability. Utilizing the robust computational capabilities of the Summit, extensive experiments were conducted on diverse datasets such as Pets, MedMNist, Flowers, and Flowers102. Our findings reveal that both overly large and overly small batch sizes can lead to performance instability and affect accuracy. However, the relationship between batch size and unlearnability varied across datasets, highlighting the necessity for tailored batch size strategies to achieve optimal data protection. The use of Summit’s high-performance GPUs, along with the efficiency of the DDP framework, facilitated rapid updates of model parameters and consistent training across nodes. Our results underscore the critical role of selecting appropriate batch sizes based on the specific characteristics of each dataset to prevent learning and ensure data security in deep learning applications. The source code is publicly available at https: // github. com/ hrlblab/ UE_ HPC .

Zhu, Yanfan [Vanderbilt University, Nashville, TN,↗

ISRU Technology Development for Extraction of Water from the Mars Surface

Goals: Develop technologies to extract water from planetary regolith considering production rate vs: energy/power consumption (efficiency, yield, heat recuperation options); Mass and sizing: modularity, batch sizes, soil feed options, etc.; Ruggedness in terms of soil, environmental, and operational parameters; seals and component wear, etc.; Removal of product and disposal of spent material. Summary: Technology development is underway for several ISRU water extraction hardware concepts for Mars application - Hydrated minerals (Auger dryer, Microwave, Open Air), Subsurface Ice (Rodwell); Models are developed with experimental and breadboard efforts for use in larger ISRU system models; Each effort consists of a 3 year development plan, with the goal of integrating into a larger subsystem test in 2020 - Concurrent technology advance allows for flexibility in system design; depending on architecture decisions and progress of associated subsystems.

In-Situ Resource Utilization↗

Tula: Optimizing Time, Cost, and Generalization in Distributed Large-Batch Training

Distributed training increases the number of batches processed per iteration either by scaling-out (adding more nodes) or scaling-up (increasing the batch-size). However, the largest configuration does not necessarily yield the best performance. Horizontal scaling introduces additional communication overhead, while vertical scaling is constrained by computation cost and device memory limits. Thus, simply increasing the batch-size leads to diminishing returns: training time and cost decrease initially but eventually plateaus, creating a knee-point in the time/cost vs. batch-size pareto curve. The optimal batch-size therefore depends on the underlying model, data and available compute resources. Large batches also suffer from worse model quality due to the well-known “generalization gap”. In this paper, we present Tula, an online service that automatically optimizes time, cost, and convergence quality for large-batch training of convolutional models. It combines parallel-systems modeling with statistical performance prediction to identify the optimal batchsize. Tula predicts training time and cost within 7.5−14% error across multiple models, and achieves up to 20× overall speedup and improves test accuracy by ≈9% on average over standard large-batch training on various vision tasks, thus successfully mitigating the generalization gap and accelerating training at the same time.

Tyagi, Sahil [ORNL] (ORCID:0009000783144745)↗

Autonomous organic synthesis for redox flow batteries via flexible batch Bayesian optimization

Traditional trial-and-error methods for materials discovery are inefficient to meet the urgent demands posed by the rapid progression of climate change. This urgency has driven the increasing interest in integrating robotics and machine learning into materials research to accelerate experimental learning. However, idealized decision-making frameworks to achieve maximum sampling efficiency are not always compatible with high-throughput experimental workflows inside a laboratory. For multi-step chemical processes, differences in hardware capacities can complicate the digital framework by introducing constraints on the maximum number of samples in each step of the experiment, hence causing varying batch sizes in variable selection within the same batch. Therefore, designing flexible sampling algorithms is necessary to accommodate the multi-step synthesis with practical constraints unique to each high-throughput workflow. In this work, we designed and employed three strategies on a high-throughput robotic platform to optimize the sulfonation reaction of redox-active molecules used in flow batteries. Our strategies adapt to the multi-step experimental workflow, where their formulation and heating steps are separate, causing varying batch size requirements. By strategically sampling using clustering and mixed-variable batch Bayesian optimization, we were able to iteratively identify optimal conditions that maximize the yields. Our work presents a flexible approach that allows tailoring the machine learning decision-making to suit the practical constraints in individual high-throughput experimental platforms, followed by performing resource-efficient yield optimization using available open-source Python libraries.

Tamura, Clara [Univ. of Washington, Seattle, WA (U↗

Active Learning‐Driven Inkless Additive Nanomanufacturing for Printed Electronics

Inkless additive nanomanufacturing for printed electronics promises broad material and substrate versatility, yet the high-dimensional print parameter space makes tuning print parameters time-intensive. We present a Bayesian optimization study that constructs a digital twin from printed-silver data to benchmark surrogate models, acquisition functions, and batch sizes head-to-head to achieve user-specified target resistance. Tested surrogate models included Gaussian process, random forest, and Bayesian neural network surrogates with expected improvement and confidence bound acquisition functions. In total, we evaluate 48 unique model configurations alongside a random sampling baseline for comparison. For printed silver, the Bayesian neural network with a batch size of one achieved the lowest average cumulative regret, approximately four times more efficient on average than random sampling. To balance performance and substrate space, a random forest model with expected improvement and a batch size of four was chosen as the model for validation testing. Applying this chosen configuration to copper with an additional print parameter, the model achieved a resistance within 0.15 Ω of a 1 Ω target in fewer than 30 printed lines across five validation sets. Altogether, the workflow yields a tuned and validated model that efficiently guides experiments toward the target while simultaneously learning the parameter space.

Bevel, Colton [Auburn University, AL (United State↗

spdlayers

Symmetric Positive Definite (SPD) enforcement layers for PyTorch. Regardless of the input, the output of these layers will always be a SPD tensor! The `Cholesky` layer uses a cholesky factorization to enforce SPD, and the `Eigen` layer uses an eigendecomposition to enforce SPD. Both layers take in some tensor of shape `[batch_size, input_shape]` and output a SPD tensor of shape`[batch_size, output_shape, output_shape]`. The relationship between input and output is defined by the following. ```python input_shape = sum([i for i in range(output_shape + 1)]) ``` The layers have no learnable parameters, and merely serve to transform a vector space to a SPD matrix space.

Jekel, CharlesF.↗

Microwave Structure Construction Capability Year One Accomplishments

The Microwave Structure Construction Capability (MSCC) element, part of the Moon to Mars Planetary Autonomous Construction Project (MMPACT) was initiated in 2020. MSCC is responsible for creating horizontal and vertical infrastructure on the moon using microwave energy. Microwave energy was selected since it is the only method to volumetrically heat the regolith. All other sintering/melting methods rely on thermal conduction through the very low conductivity surface, resulting in an inefficient process. Advances were achieved in materials characterization and understanding, microwave sintering in vacuum, and microwave design and analyses. Two dielectric property testing systems have been developed at Radiance Technologies and JPL. These will examine dielectric properties at cryogenic temperatures and over a broad frequency range. Permittivity and permeability testing at -60 ̊C in vacuum from 0.05 to 3 GHz has been generated at JPL. Additional modifications will be made to go to -190 ̊C (LN2). Radiance Technologies created a test system to measure dielectric properties at greater than 10 GHz and initiated work on developing a vacuum capable, portable test system to measure dielectric properties of Apollo regolith and simulants from 100 MHz to 18 GHz. These tests are to identify optimal heating frequencies and protocols. During microwave sintering at about 1100 ̊C, volatiles were creating difficulty in achieving a reasonably dense specimen. Due to processing in vacuum and the nature of the lunar regolith, some volatiles and porosity are expected. However, the Earth produced simulants have non-lunar materials in them that create volatiles that aren’t representative of lunar regolith. Therefore, a five month effort was conducted to establish a heat treat method to remove these non-lunar materials. Tests were conducted using TGA mass spectrometry, heating in vacuum and conducting mass spectrometry, dielectric and DTA, Raman, BET, particle size analysis, morphological analysis, carbon and sulfur chemical content determination and microscopy. The process has been scaled-up to 6 kg batch size and undergoing evaluation. A 36 kg batch size is the target for JSC-1A and other limited availability simulants. These calcining protocols will be standard for NASA and beyond. MSCC has also created scalable processes for fabricating synthetic lunar materials. Processes to fabricate Anorthite (plagioclase CaAl2Si2O8), Diopside (pyroxene CaMgSi₂O₆), and Enstatite (pyroxene Mg2Si2O6) have been generated. These materials will enable generation of microwave sintering models to bound various composition ranges anticipated on the Moon, therefore mitigating the need for a precise simulant with respect to location on the Moon. Successful microwave sintering in air using a horn applicator was demonstrated. All previous microwave vacuum sintering in the literature was at small scale and in a contained enclosure thus taking advantage of reflections. This is the first to use a lunar like microwave applicator to sinter ceramic in a bed as it would be done on the Moon. Small scale and inert sintering were conducted to assist in developing protocols with quicker turnaround times than larger scale testing. Testing has anchored thermal analysis predicting heat flow in vacuum during microwave sintering. Thermal conductivity testing was also initiated. Microwave coupling to the regolith has been modeled by multiple organizations and with different software packages. At least six horn designs and applicator configurations for both magnetron and solid state sources are being examined. Optimal simulant container designs for microwaves have also been generated. The power and electronics design for the solid state microwave system has been initiated. Concept designs for a lander based microwave sintering have been evaluated.

microwave↗

Component and System Sensitivity Considerations for Design of a Lunar ISRU Oxygen Production Plant

Component and system sensitivities of some design parameters of ISRU system components are analyzed. The differences between terrestrial and lunar excavation are discussed, and a qualitative comparison of large and small excavators is started. The effect of excavator size on the size of the ISRU plant's regolith hoppers is presented. Optimum operating conditions of both hydrogen and carbothermal reduction reactors are explored using recently developed analytical models. Design parameters such as batch size, conversion fraction, and maximum particle size are considered for a hydrogen reduction reactor while batch size, conversion fraction, number of melt zones, and methane flow rate are considered for a carbothermal reduction reactor. For both reactor types the effect of reactor operation on system energy and regolith delivery requirements is presented.

Linne, Diane L.↗

EdgeAI: Machine learning via direct attached accelerator for streaming data processing at high shot rate x-ray free-electron lasers

We present a case for low batch-size inference with the potential for adaptive training of a lean encoder model. We do so in the context of a paradigmatic example of machine learning as applied in data acquisition at high data velocity scientific user facilities such as the Linac Coherent Light Source-II x-ray Free-Electron Laser. We discuss how a low-latency inference model operating at the data acquisition edge can capitalize on the naturally stochastic nature of such sources. We simulate the method of attosecond angular streaking to produce representative results whereby simulated input data reproduce high-resolution ground truth probability distributions. By minimizing the mean-squared error between the decoded output of the latent representation and the ground truth distributions, we ensure that the encoding layers and resulting latent representation maintains full fidelity for any downstream task, be it classification or regression. We present throughput results for data-parallel inference of various batch sizes, some with throughput exceeding 100 k images per second. We also show in situ training below 10 s per epoch for the full encoder–decoder model as would be relevant for streaming and adaptive real-time data production at our nation’s scientific light sources.

97 MATHEMATICS AND COMPUTING↗

Active learning for SNAP interatomic potentials via Bayesian predictive uncertainty

Bayesian inference with a simple Gaussian error model is used to efficiently compute prediction variances for energies, forces, and stresses in the linear SNAP interatomic potential. Here, the prediction variance is shown to have a strong correlation with the absolute error over approximately 24 orders of magnitude. Using this prediction variance, an active learning algorithm is constructed to iteratively train a potential by selecting the structures with the most uncertain properties from a pool of candidate structures. The relative importance of the energy, force, and stress errors in the objective function is shown to have a strong impact upon the trajectory of their respective net error metrics when running the active learning algorithm. Batched training of different batch sizes is also tested against singular structure updates, and it is found that batches can be used to significantly reduce the number of retraining steps required with only minor impact on the active learning trajectory.

97 MATHEMATICS AND COMPUTING↗

Demonstration of the Reproducibility Challenges in the Sintering Behavior of Lithium‐Stuffed Garnets in Scaling up Synthesis

Lithium-stuffed garnets, such as Li 7 La 3 Zr 2 O 12 (LLZO), are promising candidates for next-generation solid-state batteries because of their high room-temperature ionic conductivity and chemical stability against lithium metal anodes, which are crucial for achieving higher energy density. However, realizing LLZO's potential in practical devices requires synthesis methods that can be scaled reliably to large batch sizes for manufacturing. Herein, we investigate the sintering reproducibility of LLZO synthesized at larger scales using ultrasonic spray pyrolysis, a cost-effective and scalable synthesis route. Two 100 g batches of Al-doped LLZO are prepared and their sintering behavior is examined in detail. Both Al-LLZO batches contain over 90 wt.% cubic-phase LLZO, and both batches exhibit room temperature conductivities greater than 1 × 10 −4 S cm −1 at a relative density above 0.8. However, variations in secondary phases and subtle differences in Al content lead to significant differences in densification and microstructure. These results demonstrate that LLZO's sintering behavior is highly sensitive to small changes in secondary phases and Al content, creating reproducibility challenges when moving from laboratory- to manufacturing-scale synthesis.

36 MATERIALS SCIENCE↗

Finding MIDDLE Ground: Scalable and Secure Distributed Learning

Edge computing methods allow devices to efficiently train a high-performing, robust, and personalized model for predictive tasks. However, these methods succumb to privacy and scalability concerns such as adversarial data recovery and expensive model communication. Furthermore, edge computing methods unrealistically assume that all devices train an identical model. In practice, edge devices have varying computational and memory constraints which may not allow certain devices to have the space or speed to train a specific model. To overcome these issues, we propose MIDDLE: a model independent distributed learning algorithm which allows heterogeneous edge devices to assist each other’s training while communicating only non-sensitive information. MIDDLE unlocks the ability for edge devices, regardless of computational or memory constraints, to assist each other even with completely different model architectures. Furthermore, MIDDLE does not require model or gradient communication which greatly reduces communication size and time. We prove that MIDDLE attains the optimal convergence rate O(1/sqrt(TM)) of stochastic gradient descent for convex and non-convex smooth optimization (for total iterations T and batch size M). Finally, our experimental results demonstrate that MIDDLE (even in non-IID data settings) attains robust and high-performing models without model or gradient communication.

Bornstein, Marc I.↗

Variance-Reduced Accelerated First-Order Methods: Central Limit Theorems and Confidence Statements

In this paper, we consider a strongly convex stochastic optimization problem and propose three classes of variable sample-size stochastic first-order methods: (i) the standard stochastic gradient descent method, (ii) its accelerated variant, and (iii) the stochastic heavy-ball method. In each scheme, the exact gradients are approximated by averaging across an increasing batch size of sampled gradients. We prove that when the sample size increases at a geometric rate, the generated estimates converge in mean to the optimal solution at an analogous geometric rate for schemes (i)–(iii). Based on this result, we provide central limit statements, whereby it is shown that the rescaled estimation errors converge in distribution to a normal distribution with the associated covariance matrix dependent on the Hessian matrix, the covariance of the gradient noise, and the step length. If the sample size increases at a polynomial rate, we show that the estimation errors decay at a corresponding polynomial rate and establish the associated central limit theorems (CLTs). Under certain conditions, we discuss how both the algorithms and the associated limit theorems may be extended to constrained and nonsmooth regimes. As a result, we provide an avenue to construct confidence regions for the optimal solution based on the established CLTs and test the theoretical findings on a stochastic parameter estimation problem.

Lei, Jinlong↗

Custom Equipment Development for Processing of Surplus Plutonium

The Strategic Laboratory Assessment (SLA), a collaborative team of SRNL and ORNL personnel, has been established to advance the objectives of the Surplus Plutonium Disposition (SPD) Project, by identifying and developing technologies to accelerate disposition, reduce life cycle costs, minimize worker radiation exposure, improve worker safety, and minimize Surplus Plutonium Disposition Program risks. [1] The SLA team has identified can cutting and plutonium (Pu) oxide size reduction as two glovebox processes where technology enhancements would be valuable. The DOESTD-3013 package currently in use for Pu downblending requires cutting two nested cans before the inner convenience can that holds the Pu oxide may be accessed for further processing. A rotary tubing-style cutter is used for opening the 3013 packages within the glovebox. Collet changeouts are required between cutting of the outer and inner cans. The SLA team is currently developing and testing an adjustable-clamp can cutter design that eliminates collet changeouts and allows cutting of the outer and inner can at the same time, resulting in significant reduction of radiological dose and process time, as well as improved ergonomics. To meet the Pu oxide particle size requirement, size reduction of Pu oxide agglomerations must be performed within the process gloveboxes. The SLA team has identified jaw crushing technology as an alternative to the currently employed rotary mill. Jaw crusher advantages include reduced dust within the glovebox, increased batch sizes, and easier integration with other glovebox processes due to the flow-through nature of jaw crushing. Commercially manufactured jaw crushers are either too large and/or too heavy for implementation in the SPD gloveboxes, so the SLA team is developing and testing a custom jaw crusher to meet the needs of the SPD Project.

Krementz, Daniel [Savannah River National Laborato↗

An Accelerated Clip Algorithm for Unstructured Meshes: A Batch-Driven Approach

The clip technique is a popular method for visualizing complex structures and phenomena within 3D unstructured meshes. Meshes can be clipped by specifying a scalar isovalue to produce an output unstructured mesh with its external surface as the isovalue. Similar to isocontouring, the clipping process relies on scalar data associated with the mesh points, including scalar data generated by implicit functions such as planes, boxes, and spheres, which facilitates the visualization of results interior to the grid. In this paper, we introduce a novel batch-driven parallel algorithm based on a sequential clip algorithm designed for high-quality results in partial volume extraction. Our algorithm comprises five passes, each progressively processing data to generate the resulting clipped unstructured mesh. The novelty lies in the use of fixed-size batches of points and cells, which enable rapid workload trimming and parallel processing, leading to a significantly improved memory footprint and run-time performance compared to the original version. On a 32-core CPU, the proposed batch-driven parallel algorithm demonstrates a run-time speed-up of up to 32.6x and a memory footprint reduction of up to 4.37x compared to the existing sequential algorithm. The software is currently available under an open-source license in the VTK visualization system.

Tsalikis, Spiros↗

Latency Analysis of the Nexus Digital Twin Framework

Real-time digital catalogs are increasingly relied upon to track metadata and connect disparate data sources for cloud-based data integration efforts. One such tool, Deeplynx Nexus is supporting real-time digital twin efforts through event-driven data integration and time-series queries. Nexus’s usefulness for these applications depends critically on how quickly individual records can be uploaded and downloaded, since delays directly affect the responsiveness of any system built on top of it. However, the actual latency a user should expect from Nexus has not been systematically measured before, particularly for the small, frequent transactions typical of live sensor feeds. Here we show that single-record round-trip latency is 61.1 ms on a local Nexus instance and 391.7 ms on the hosted production infrastructure, a roughly 6.4x difference driven primarily by fixed per-request overhead rather than data volume. This overhead dominates at small scale: comparing single-record and ten-record trials suggests approximately 56 ms of each single-record request is fixed connection and authentication cost rather than data-transfer time, meaning batching even a handful of records is substantially more efficient than transmitting them individually. At large batch sizes, this pattern reverses for uploads, which converge to near parity between local and hosted environments by 25,000-50,000 records, while download latency remains persistently 5.7-6.4x slower on hosted infrastructure even at scale. These results suggest that Nexus deployments intended for real-time digital twin applications should prioritize record batching over single-record transactions, and that download-path optimization on hosted infrastructure offers the largest remaining opportunity to reduce latency at scale. We anticipate these baseline measurements will serve as a reference point for future digital twin projects evaluating whether Nexus’s latency profile meets their real-time requirements, and as a benchmark for tracking the effect of future infrastructure or API changes.

99 - GENERAL AND MISCELLANEOUS↗