Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “batch size”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Batched sparse direct solver design and evaluation in SuperLU_DIST

Over the course of interactions with various application teams, the need for batched sparse linear algebra functions has emerged in order to make more efficient use of the GPUs for many small and sparse linear algebra problems. In this paper, we present our recent work on a batched sparse direct solver for GPUs. The sparse LU factorization is computed by the levels of the elimination tree, leveraging the batched dense operations at each level and a new batched Scatter GPU kernel. The sparse triangular solve is computed by the level sets of the directed acyclic graph (DAG) of the triangular matrix. Batched operations overcome the large overhead associated with launching many small kernels. For medium sized matrix batches with not-so-small bandwidth, using an NVIDIA A100 GPU, our new batched sparse direct solver is orders of magnitude faster than a batched banded solver and uses less than one-tenth of the memory.

Boukaram, Wajih↗

BUTTER - Empirical Deep Learning Dataset

The BUTTER Empirical Deep Learning Dataset represents an empirical study of the deep learning phenomena on dense fully connected networks, scanning across thirteen datasets, eight network shapes, fourteen depths, twenty-three network sizes (number of trainable parameters), four learning rates, six minibatch sizes, four levels of label noise, and fourteen levels of L1 and L2 regularization each. Multiple repetitions (typically 30, sometimes 10) of each combination of hyperparameters were preformed, and statistics including training and test loss (using a 80% / 20% shuffled train-test split) are recorded at the end of each training epoch. In total, this dataset covers 178 thousand distinct hyperparameter settings ("experiments"), 3.55 million individual training runs (an average of 20 repetitions of each experiments), and a total of 13.3 billion training epochs (three thousand epochs were covered by most runs). Accumulating this dataset consumed 5,448.4 CPU core-years, 17.8 GPU-years, and 111.2 node-years.

Array↗

Effect of particle size on the capture of uranium oxide colloidal particles from aqueous suspensions via high-gradient magnetic filtration

The effectiveness of High Gradient Magnetic Filtration (HGMF) in capturing uranium oxide particles from suspensions was investigated in this study. Two sets of experiments were performed to evaluate the importance of size on the capture of uranium oxide particles. The first considered two batches sieved into size bins of< 5, 5–10, 10–15, and 15–20 µm, while the second was performed using two suspensions with diameters smaller than 1.0 µm and between 1.0 and 1.5 µm. Iron oxide experiments, with particles between 0.3 and 0.8 µm, were performed for calibration purposes. In all experiments, a surfactant (Triton-X100 or sodium dodecyl sulfate) was used to prevent particle aggregation and limit the influence of non-magnetic capture mechanisms. A magnetic field of approximately 1.1 Tesla was generated using a water cooled electromagnet. HGMF was performed using tubular filters packed with ferromagnetic stainless-steel wool. Of the initial four uranium oxide particle sizes, magnetic capture was only observed for particles with a diameter of less than 5 µm, while larger particles experienced no magnetic and minimal total capture. For particles with diameters smaller than 1.0 µm and between 1.0 and 1.5 µm, capture efficiencies increased by 39 ± 9% and 34 ± 6% respectively, solely due to the magnetic field. Although the magnetic force is proportional to particle diameter, the capture efficiency decreased as diameter increased. So these results suggest that Brownian diffusion, which is influential for micron sized particles and increases with decreasing particle size, is acting in conjunction with the magnetic force to influence the efficacy of HGMF for uranium oxide. This important finding underscores the effectiveness of Brownian diffusion in increasing the rate of collision between particles and collector fibers. A stochastic trajectory model was developed to incorporate the influence of Brownian motion on particle behavior and filter removal efficiency. Modeling results are discussed and compared for uranium and iron oxide particles.

42 ENGINEERING↗

Performance Assessment of AFQMC Implementation (ECP Milestone 5.2 Report)

The auxiliary field quantum Monte Carlo algorithm in QMCPACK was ported to run on AMD GPUs using HIP. Initial performance measurements on early access hardware indicate a 4x slowdown compared to NVIDIA V100s. The slowdown was determined to be a result of suboptimal ROCM libraries for small matrix sizes coupled with some poorly performing hipified CUDA kernels. Through a preliminary analysis of HIP kernel performance we determined a strategy for pro ling and tuning kernels. We request that ROCM libraries are optimized for smaller problem sizes particularly for batched operations.

97 MATHEMATICS AND COMPUTING↗

Deep Generative Models that Solve PDEs: Distributed Computing for Training Large Data-Free Models

Recent progress in scientific machine learning (SciML) has opened up the possibility of training novel neural network architectures that solve complex partial differential equations (PDEs). Several (nearly data free) approaches have been recently reported that successfully solve PDEs, with examples including deep feed forward networks, generative networks, and deep encoder-decoder networks. However, practical adoption of these approaches is limited by the difficulty in training these models, especially to make predictions at large output resolutions (≥1024×1024). Here we report on a software framework for data parallel distributed deep learning that resolves the twin challenges of training these large SciML models - training in reasonable time as well as distributing the storage requirements. Our framework provides several out of the box functionality including (a) loss integrity independent of number of processes, (b) synchronized batch normalization, and (c) distributed higher-order optimization methods. We show excellent scalability of this framework on both cloud as well as HPC clusters, and report on the interplay between bandwidth, network topology and bare metal vs cloud. We deploy this approach to train generative models of sizes hitherto not possible, showing that neural PDE solvers can be viably trained for practical applications. We also demonstrate that distributed higher-order optimization methods are 2-3× faster than stochastic gradient-based methods and provide minimal convergence drift with higher batch-size.

PDEs↗

Acceleration of Graph Neural Network-Based Prediction Models in Chemistry via Co-Design Optimization on Intelligence Processing Units

Atomic structure prediction and associated property calculations are the bedrock of chemical physics. Since high-fidelity ab initio modeling techniques for computing the structure and properties can be prohibitively expensive, this motivates the development of machine-learning (ML) models that make these predictions more efficiently. Training graph neural networks over large atomistic databases introduces unique computational challenges such as the need to process millions of small graphs with variable size and support communication patterns that are distinct from learning over large graphs such as social networks. We demonstrate a novel hardware-software co-design approach to scale up the training of atomistic graph neural networks (GNN) for structure and property prediction. First, to eliminate redundant computation and memory associated with alternative padding techniques and to improve throughput via minimizing communication, we formulate the effective coalescing of the batches of variable-size atomistic graphs as the bin packing problem and introduce a hardware-agnostic algorithm to pack these batches. In addition, we propose hardware-specific optimizations including a planner and vectorization for the gather-scatter operations targeted for Graphcore’s Intelligence Processing Unit (IPU), as well as model-specific optimizations such as merged communication collectives and optimized softplus. Putting these all together, we demonstrate the effectiveness of the proposed co-design approach by providing an implementation of a well-established atomistic GNN on the Graphcore IPUs. We evaluate the training performance on multiple atomistic graph databases with varying degrees of graph counts, sizes and sparsity. Here, we demonstrate that such a co-design approach can reduce the training time of atomistic GNNs and can improve the performance by up to 1.5× compared to the baseline implementation of the model on the IPUs. Additionally, we compare our IPU implementation with a Nvidia GPU-based implementation and show that our atomistic GNN implementation on the IPUs can run 1.8× faster on average compared to the execution time on the GPUs.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Variations in GARS powder microstructure as a function of powder chemistry and particle size

The properties of metal produced through powder metallurgy depends on the feedstock used. Powders produced via gas atomization reaction synthesis (GARS) are used to produce oxide dispersion strengthened alloys. The desired powder size range can vary for each consolidation technique. However, powder microstructure also can vary with powder particle size, which in turn can impact the microstructure and properties of the consolidated parts. In this study, GARS powders are characterized via inductively coupled plasma mass spectroscopy, inert gas fusion, and high-resolution x-ray diffraction to determine variations in elemental and phase compositions. Transmission electron microscopy was used to understand microstructure variations as a function of chemistry and size. Across the three batches tested intermetallic content was 0.73–1.35 wt% in the 0-20 μm powder batch and increased to 2.46–3.80 wt% in the coarse 45-106 μm batch. Across all batches, volume percent of surface oxidation decreased with powder diameter, with volume percents within the range of 0.75–1.2 % across 10 μm powder particles, and below 0.4 % across coarse powder particles approximately 100 μm in diameter. These observations were supported by inert gas fusion measurements. However, the oxide layer was thicker in coarse powder particles due to a slower cooling rate. Increasing oxygen content in atomization gas to 2000 ppm and adding yttrium increased both the surface oxidation content and yttrium intermetallic content. Lastly, intermetallic phases within the powder coarsened with powder size. Intermetallic morphology changed from fine spherical intermetallic and columnar dendritic growth to a cellular structure with finer spherical intermetallic, to coarse irregular intermetallic and intermetallic along grain boundaries as a result of slower cooling rate and solidification rate in coarse powder particles. Furthermore, the addition of zirconium does not appear to significantly change intermetallic morphology, but the composition changed from a Y-Fe rich intermetallic to a Y-Zr-Fe intermetallic.

42 ENGINEERING↗

A physics-informed and hierarchically regularized data-driven model for predicting fluid flow through porous media

This paper presents a new deep learning data-driven model for predicting structure dependent pore-fluid velocity fields in rock. The model is based on a Convolutional Auto-Encoder (CAE) artificial neural network capable of learning from image data generated by direct numerical simulations of fluid flow through pore-structures, such as by Lattice Boltzmann or molecular dynamics methods. The main novelty of the model in comparison to previous CAE-based data-driven approaches consists of three parts. The first is a methodology for decomposing the full-domain of the porous media into sub-regions, or “sub-domains”, in order to reduce the overall size of the CAE, batch process the sub-domains in parallel, and enable the CAE to learn local and generalizable nonlinear mappings of pore-fluid velocities. The second consists of embedding the finite difference solutions of the incompressible Navier-Stokes and continuity equations into convolutional layers prior to the CAE in order to provide the CAE with knowledge of fluid dynamics physics (PhyFlow). The third main novelty is that the training of the CAE is regularized with a hierarchical loss function that encourages the learning of fluid flow patterns (in a way similar to ranked modes in principal component analysis), ranking from most to least important. This is shown to increase the stability in learning, reduce over-fitting, and promote interpretability of the CAE neural network layers (HierCAE). The comprehensive new data-driven model, which we call the PhyFlow-HierCAE model, is shown to exhibit improved accuracy and generalizability of flow field predictions over conventional CAE models, attributable to the embedded physical knowledge and the hierarchical regularization, as well as realize orders of magnitude speed-ups in computation times as a surrogate for the direct numerical simulations. Examples of training and forward predictions on unseen pore-structures are provided and evaluated for data from Lattice Boltzmann and molecular dynamics simulations of pore-fluid flow. The model is shown to be a fast and accurate emulator (or “surrogate”) for predicting effective permeability of unseen pore-structures based on learning from relatively small direct numerical simulation datasets.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Online Parameter Estimation Methods for Adaptive Cruise Control Systems

Modeling Adaptive Cruise Control (ACC) vehicles enables the understanding of the impact of these vehicles on traffic flow. In this work, two online methods are used to provide real time system identification of ACC enabled vehicles. The first technique is a recursive least squares (RLS) approach, while the second method solves a nonlinear joint state and parameter estimation problem via particle filtering (PF). We provide a parameter identifiability analysis for both methods to analytically show that the model parameters are not identifiable using equilibrium driving. The accuracy and computational runtime of the online methods are compared to a commonly used offline simulation-based optimization (i.e., batch optimization) approach. The methods are tested on synthetic data as well as on empirical data collected directly from a 2019 model year ACC vehicle using data from sensors that are part of the stock ACC system. The online methods are scalable and provide comparable accuracy to the batch method. RLS runs in real time and is two orders of magnitude faster than the batch method for modest sized (e.g., 15 min) datasets. The particle filter also runs in real- time, and is also suitable in streaming applications in which the datasets can grow arbitrarily large.

33 ADVANCED PROPULSION SYSTEMS↗

Serial femtosecond and serial synchrotron crystallography can yield data of equivalent quality: A systematic comparison

For the two proteins myoglobin and fluoroacetate dehalogenase, we present a systematic comparison of crystallographic diffraction data collected by serial femtosecond (SFX) and serial synchrotron crystallography (SSX). To maximize comparability, we used the same batch of micron-sized crystals, the same sample delivery device, and the same data analysis software. Overall figures of merit indicate that the data of both radiation sources are of equivalent quality. For both proteins, reasonable data statistics can be obtained with approximately 5000 room-temperature diffraction images irrespective of the radiation source. The direct comparability of SSX and SFX data indicates that the quality of diffraction data obtained from these samples is linked to the properties of the crystals rather than to the radiation source. Therefore, for other systems with similar properties, time-resolved experiments can be conducted at the radiation source that best matches the desired time resolution.

59 BASIC BIOLOGICAL SCIENCES↗

Carbon nanospike coated nanoelectrodes for measurements of neurotransmitters

Here, carbon nanoelectrodes enable the detection of neurotransmitters at the level of single cells, vesicles, synapses and small brain structures. Previously, the etching of carbon fibers and 3D printing based on direct laser writing have been used to fabricate carbon nanoelectrodes, but these methods lack the ability of mass manufacturing. In this paper, we mass fabricate carbon nanoelectrodes by growing carbon nanospikes (CNSs) on metal wires. CNSs have a short, dense and defect-rich surface that produces remarkable electrochemical properties, and they can be mass fabricated on almost any substrate without using catalysts. Tungsten wires and niobium wires were electrochemically etched in batch to form sub micrometer sized tips, and a layer of CNSs was grown on the metal wires using plasma-enhanced chemical vapor deposition (PE-CVD). The thickness of the CNS layer was controlled by the deposition time, and a thin layer of CNSs can effectively cover the entire metal surface while maintaining the tip size within the sub micrometer scale. The etched tungsten wires produced tapered conical nanotips, while the etched niobium wires were long and thin. Both showed excellent sensitivity for the detection of outer sphere ruthenium hexamine and the inner sphere test compound ferricyanide. The CNS nanosensors were used for the measurement of dopamine, serotonin, ascorbic acid and DOPAC with fast-scan cyclic voltammetry. The CNS nanoelectrodes had a large surface area and numerous defect sites, which improved the sensitivity, electron transfer kinetics and adsorption. Finally, the CNS nanoelectrodes were compared with other nanoelectrode fabrication methods, including flame etching, 3D printing, and nanopipettes, which are slower to make and more difficult for mass fabrication. Thus, CNS nanoelectrodes are a promising strategy for the mass fabrication of nanoelectrode sensors for neurotransmitters.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Pore connectivity influences mass transport in natural rocks: Pore structure, gas diffusion and batch sorption studies

For this work, six rocks (one granodiorite, one limestone, two chalks, one mudstone, and one dolostone) with different extents of heterogeneity at six different particle sizes (from 75 to 8000 μm) were studied to describe the effects of pore connectivity on mass transport. The methods applied were (i) porosity measurement of granular rocks, (ii) analyses of gas-phase diffusive transport in a bed of packed particles, along with a solid quartz method at these six particle sizes being developed to identify the contribution of intraparticle diffusion, and (iii) batch sorption tests of multiple ions (anions and cations) with subsequent analyses of inductively coupled plasma-mass spectrometry. Granular porosity measurement results reveal that with decreasing particle sizes, the effective porosities for the “heterogenous” group of rocks (Grimsel granodiorite and Edwards limestone) increase, whereas the porosities of another “homogeneous” group (two Israel chalk samples, Japan mudstone, and Wyoming dolostone) remain constant. Gas diffusion results show that the intraparticle gas diffusion coefficient among these two sample groups, varying in the magnitude of 10 -8 to 10 -6 m 2 /s, are not directly correlated to the porosity differences. Moreover, the batch sorption work displays a different affinity of rocks for various tracers. For Grimsel granodiorite, Japan mudstone, and Wyoming dolostone, the adsorption capacity of Sm 3+ and Eu 3+ increases as the particle size decreases. In general, this integrated research of grain size distribution, granular rock porosity, intraparticle diffusivity, and ionic sorption capacity gives insights into the pore connectivity effect on both physical and chemical transport behaviors for different lithologies and/or different particle sizes.

58 GEOSCIENCES↗

Amino acids as performance-controlling additives in carbonation-activated cementitious materials

This article presents an investigation on the application of amino acids to control the CaCO{sub 3} crystallization in carbonation cured wollastonite composites. It was observed that wollastonite carbonated without any amino acid formed calcite as the primary polymorph of CaCO{sub 3}. In contrast, the use of amino acids as admixtures resulted in the formation of stable amorphous calcium carbonate (ACC), vaterite, and aragonite during the carbonation of wollastonite. The carbonated composites produced with amino acids were observed to have a lower critical pore size, but a higher total porosity, compared to the control batch. Additionally, the utilization of amino acids was observed to increase the flexural strength and compressive strength of the composites up to 106% and 48%, respectively, compared to the control batch. Such performance enhancement of the carbonated composites in the presence of amino acids was attributed to the reduced critical pore size and the formation of organic-inorganic hybrid phases in the matrix.

36 MATERIALS SCIENCE↗

Scalable Synthesis of Pt/SrTiO 3 Hydrogenolysis Catalysts in Pursuit of Manufacturing-Relevant Waste Plastic Solutions

Here, an improved hydrothermal synthesis for shape-controlled, size-controlled 60 nm SrTiO 3 nanocuboid (STO NC) supports, which facilitates the scalable creation of platinum nanoparticles catalyst supported on STO (Pt/STO) for the chemical conversion of waste polyolefins, is reported herein. This synthetic method: 1) produces STO NC supports with average sizes ranging from 25 – 80 nm with narrow size distributions 2) demonstrates how SrCO 3 formation and variation in solution pH prevent the formation of STO NCs, and 3) establishes that STO nucleation prior to the hydrothermal treatment favors nanocuboid formation. The updated hydrothermal synthesis was scaled-up and conducted in a 4L batch reactor, resulting in STO NCs of comparable size and morphology (m = 22.5 g, d avg = 58.6 ± 16.2 nm) to those synthesized under standard hydrothermal conditions in a lab-scale 125 mL autoclave reactor. Size-controlled STO NCs, ranging in roughly 10 nm increments from the 25 nm to 80 nm, were used to support Pt deposited through strong electrostatic adsorption (SEA), a practical and scalable solution-based method. Using SEA techniques and a STO support with an average size of 39.3 ± 6.3 nm, a Pt/STO catalyst with 3.6 wt% Pt was produced and used for high-density polyethylene hydrogenolysis under previously-reported conditions (170 psi H 2 , 300°C, 96h; final product: M w = 2400, Ð = 1.03). As a well-established model system for studying the behavior of heterogeneous catalysts and their supports in reactions, the Pt/STO system detailed in this work presents a unique opportunity to simultaneously convert waste plastic into commercially-viable products while gaining fundamental insight into the mechanism of polyethylene hydrogenolysis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Chemometrics and visible diffuse reflectance spectroscopy to classify plutonium dioxide

Diffuse reflectance (DR) spectra in the Vis-NIR (∼380–1050 nm) region were acquired for a series of PuO 2 samples with a spot size of about 10 × 10 μm. Two batches of six PuO 2 samples, synthesized approximately 7.5 months apart, were prepared using both Pu(III) and Pu(IV) oxalate precursors at three distinct calcination temperatures (450, 650, and 950 °C). This yielded a total of 12 PuO 2 samples and 433 DR spectra. The DR spectrum of PuO 2 contained numerous peaks in the visible region, and characteristic features were identified with respect to calcination temperature and chemistry. A distinct peak multiplet near 615 nm was observed for samples prepared at low calcination temperatures, and a peak near 660 nm was observed for higher calcination temperatures. A multivariate classification strategy based on principal component analysis (PCA) was developed to distinguish PuO 2 calcination temperatures of 450, 650, and 950 °C with 100 % accuracy. Classification results also indicate the potential to distinguish chemical processing history (i.e., Pu(III) or Pu(IV)) based on the spectra with 72 % accuracy based on k-nearest neighbors applied to the PCA scores. Partial least squares discriminant analysis was used to identify variation among batches with 88 % accuracy and found that peaks near 669, 681, 811, and 970 nm were the most useful for predicting the batch identity. Here, this work demonstrates how micro-diffuse reflectance spectroscopy and chemometrics can be used to classify PuO 2 processing history based on Vis-NIR spectral features. Combining the chemometric approach with mapping sequences could provide a rapid, nondestructive approach to classify Pu oxide materials for environmental, forensics, and nonproliferation applications.

Actinide↗

Spherical powders: Control over the size and morphology of powders for additive manufacturing and enriched stable isotope nuclear targets

Metal powders are a fundamental starting point for fabricating many types of nuclear targets. Elemental powder properties can differ drastically between batches, even when using the same method. Therefore, the variation in morphology and the size of metal powders can cause variable quality and produce inconsistent results with what are otherwise proven target manufacturing techniques. Additive manufacturing has additional requirements for higher quality and more uniform feedstock. The production of spheroidized powders with uniform, reproducible properties and a narrow size distribution represents unexplored opportunities for experiments. These opportunities include experimenting with solid metals that can now flow like liquids, new options for powder handling and dispensing, and new target fabrication methods using additive manufacturing. The Stable Isotope Materials and Chemistry Group at Oak Ridge National Laboratory obtained an AMAZEMET rePowder ultrasonic metal atomization tool for creating limited batches of fully dense, free flowing, spherical powders with a narrow size distribution of extremely rare materials. Early results are presented with materials that were produced. The team explores the anticipated limits of this instrument with extremely rare materials (e.g., enriched stable isotopes) and highlights research into new fabrication techniques that provide additional options benefitting the international nuclear target community.

Zach, Mike↗

Micrometer-sized Magnetite Synthesis using Fe(OH)2(s) as a Precursor for Technetium Sequestration from Liquid Nuclear Waste Streams

Systematic batch experiments under variable adjusted physicochemical conditions were conducted to explore optimization of micrometer-sized magnetite synthesis for Tc sequestration from radionuclide waste streams using Fe(OH)2(s) as the precursor. Extensive solid characterization using x-ray diffraction and spectroscopic methods was performed to assess changes in particle morphology and size distribution, as well as Tc speciation and incorporation, in the produced mineral phases. The results show that the solution pH, temperature, and oxidation kinetics play key roles in the final mineral products. Micrometer-sized magnetite crystals (0.62-0.96 µm on average) with well-defined dodecahedral or octahedral structures were synthesized under near neutral (~pH 8) or alkaline (~pH13) conditions at 75 °C, respectively; whereas goethite dominated the end products at room temperature. An increase in pH at 75 °C improved Tc removal from 27% (near neutral pH) to 42% (alkaline pH), but the removal process remained inhibited by redox competitive Cr(VI) present in the waste streams. By adding additional Fe(II) to the system, Tc sequestration was dramatically improved to up to 87% without observable changes in the solid product. The sequestrated Tc existed as TcO2·2H2O and/or Tc(IV) incorporated into magnetite, where extended X-ray absorption fine structure (EXAFS) spectroscopy showed that more Tc was incorporated into magnetite at elevated temperatures and pH conditions, with complete Tc(IV) incorporation into magnetite occurring under 75 °C-pH 13 conditions. Our results indicate that optimal micrometer-sized magnetite can be produced for Tc sequestration by reacting Fe(OH)2(s) with a waste stream simulant under elevated pH (~13) and temperature (75 °C) conditions. The incorporation of reduced Tc(IV) into stable micrometer-sized magnetite provides a viable supplemental immobilizing technology that may be used to improve nuclear waste treatment and disposal needs.

Wang, Guohui↗