Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “learning rate”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

A phase transition for finding needles in nonlinear haystacks with LASSO artificial neural networks

To fit sparse linear associations, a LASSO sparsity inducing penalty with a single hyperparameter provably allows to recover the important features (needles) with high probability in certain regimes even if the sample size is smaller than the dimension of the input vector (haystack). More recently learners known as artificial neural networks (ANN) have shown great successes in many machine learning tasks, in particular fitting nonlinear associations. Small learning rate, stochastic gradient descent algorithm and large training set help to cope with the explosion in the number of parameters present in deep neural networks. Yet few ANN learners have been developed and studied to find needles in nonlinear haystacks. Driven by a single hyperparameter, our ANN learner, like for sparse linear associations, exhibits a phase transition in the probability of retrieving the needles, which we do not observe with other ANN learners. To select our penalty parameter, we generalize the universal threshold of Donoho and Johnstone (Biometrika 81(3):425–455, 1994) which is a better rule than the conservative (too many false detections) and expensive cross-validation. In the spirit of simulated annealing, we propose a warm-start sparsity inducing algorithm to solve the high-dimensional, non-convex and non-differentiable optimization problem. We perform simulated and real data Monte Carlo experiments to quantify the effectiveness of our approach.

97 MATHEMATICS AND COMPUTING↗

A Kaczmarz-inspired approach to accelerate the optimization of neural network wavefunctions

Neural network wavefunctions optimized using the variational Monte Carlo method have been shown to produce highly accurate results for the electronic structure of atoms and small molecules, but the high cost of optimizing such wavefunctions prevents their application to larger systems. We propose the Subsampled Projected-Increment Natural Gradient Descent (SPRING) optimizer to reduce this bottleneck. SPRING combines ideas from the recently introduced minimum-step stochastic reconfiguration optimizer (MinSR) and the classical randomized Kaczmarz method for solving linear least-squares problems. We demonstrate that SPRING outperforms both MinSR and the popular Kronecker-Factored Approximate Curvature method (KFAC) across a number of small atoms and molecules, given that the learning rates of all methods are optimally tuned. For example, on the oxygen atom, SPRING attains chemical accuracy after forty thousand training iterations, whereas both MinSR and KFAC fail to do so even after one hundred thousand iterations.

97 MATHEMATICS AND COMPUTING↗

Insights from Initial Engineering Designs of Point Source Capture at Industrial Facilities

Initial engineering design studies examining the application of state-of-the-art carbon capture technology at industrial plants contain generally overlooked real-world design considerations for near-term deployment of point source capture (PSC). The implementation of PSC across a wide range of industrial applications presents unique challenges associated with fluctuating CO<sub>2</sub> concentrations, flue gas composition, and utility and land availability. In this article, seven recent industrial retrofit PSC projects are reviewed to investigate the impact of site-specific factors on project design and cost. Across the seven projects, three capture technology classes and four industrial applications are considered, allowing insight into industry-specific opportunities for PSC technology synergy. Common challenges across projects and proposed design solutions are highlighted to propagate ideas and solutions to close the technology gaps and accelerate learning rates.

FEED Studies↗

Critically assessing sodium-ion technology roadmaps and scenarios for techno-economic competitiveness against lithium-ion batteries

Sodium-ion batteries have garnered notable attention as a potentially low-cost alternative to lithium-ion batteries, which have experienced supply shortages and price volatility for key minerals. Here we assess their techno-economic competitiveness against incumbent lithium-ion batteries using a modelling framework incorporating componential learning curves constrained by minerals prices and engineering design floors. We compare projected sodium-ion and lithium-ion price trends across over 6,000 scenarios while varying Na-ion technology development roadmaps, supply chain scenarios, market penetration and learning rates. Assuming that substantial progress can be made along technology roadmaps via targeted research and development, we identify several sodium-ion pathways that might reach cost-competitiveness with low-cost lithium-ion variants in the 2030s. In addition, we show that timelines are highly sensitive to movements in critical minerals supply chains—namely that of lithium, graphite and nickel. Our modelled outcomes suggest that being price advantageous against low-cost lithium-ion variants in the near term is challenging and increasing sodium-ion energy densities to decrease materials intensity is among the most impactful ways to improve competitiveness.

25 ENERGY STORAGE↗

Quantifying local and global mass balance errors in physics-informed neural networks

Physics-informed neural networks (PINN) have recently become attractive for solving partial differential equations (PDEs) that describe physics laws. By including PDE-based loss functions, physics laws such as mass balance are enforced softly in PINN. This paper investigates how mass balance constraints are satisfied when PINN is used to solve the resulting PDEs. We investigate PINN’s ability to solve the 1D saturated groundwater flow equations (diffusion equations) for homogeneous and heterogeneous media and evaluate the local and global mass balance errors. We compare the obtained PINN’s solution and associated mass balance errors against a two-point finite volume numerical method and the corresponding analytical solution. We also evaluate the accuracy of PINN in solving the 1D saturated groundwater flow equation with and without incorporating hydraulic heads as training data. We demonstrate that PINN’s local and global mass balance errors are significant compared to the finite volume approach. Tuning the PINN’s hyperparameters, such as the number of collocation points, training data, hidden layers, nodes, epochs, and learning rate, did not improve the solution accuracy or the mass balance errors compared to the finite volume solution. Mass balance errors could considerably challenge the utility of PINN in applications where ensuring compliance with physical and mathematical properties is crucial.

54 ENVIRONMENTAL SCIENCES↗

ZeoNet: 3D convolutional neural networks for predicting adsorption in nanoporous zeolites

Zeolites are one of the most widely used materials in the chemical industry due to their nanometer-sized pores that can adsorb and react upon molecules selectively. With hundreds of known framework topologies and hundreds of thousands of computationally predicted structures, the ability to rapidly predict zeolite performance allows researchers to prioritize their efforts on the most promising structures for a given application. Although the accuracy of forcefield-based atomistic simulations has advanced significantly in the past two decades, these simulations can be computationally expensive, especially for long-chain, complex molecules. Here, we present ZeoNet, a representation learning framework using convolutional neural networks (ConvNets) and 3D volumetric representations for predicting adsorption in zeolites. ZeoNet was trained on the task of predicting Henry's constants for adsorption, k H , of n-octadecane in more than 330 000 known and predicted zeolite materials. Employing a 3D grid based on the distances to solvent-accessible surfaces, a volumetric representation that can be generated efficiently, the best-performing ZeoNet achieved a correlation coefficient r 2 = 0.977 and a mean-squared error MSE = 3.8 in ln k H , which corresponds to an error of 9.3 kJ mol -1 in adsorption free energy. In comparison, a model based on hand-designed geometric features has values of r 2 = 0.783 and MSE = 35.7. ZeoNet is also relatively efficient and can process ≈8 structures per second on an Nvidia RTX 2080TI GPU, orders of magnitude faster than forcefield-based simulations. A systematic analysis was conducted to investigate how the choice of ConvNet architectures, the linear dimension (L) and spatial resolution (Δd) of the distance grids, batch size, optimizer, and learning rate impact the model performance. We found that ConvNets based on the ResNet architecture offer the best tradeoff between expressiveness and efficiency. The performance for all models reaches a plateau at L = 30–45 Å and depends less sensitively on grid resolution, with a small benefit around Δd = 0.30–0.45 Å. Finally, saliency maps were visualized to identify which regions of the materials contributed the most to model predictions. It was found, interestingly, that the predictions are driven primarily by the accessible pore volume rather than the region occupied by the framework atoms.

36 MATERIALS SCIENCE↗

Machine learning predictions of high-Curie-temperature materials

Technologies that function at room temperature often require magnets with a high Curie temperature, $T$ C , and can be improved with better materials. Discovering magnetic materials with a substantial $T$ C is challenging because of the large number of candidates and the cost of fabricating and testing them. Using the two largest known datasets of experimental Curie temperatures, we develop machine-learning models to make rapid $T$ C predictions solely based on the chemical composition of a material. We train a random-forest model and a k -NN one and predict on an initial dataset of over 2500 materials and then validate the model on a new dataset containing over 3000 entries. The accuracy is compared for multiple compounds' representations (“descriptors”) and regression approaches. A random-forest model provides the most accurate predictions and is not improved by dimensionality reduction or by using more complex descriptors based on atomic properties. Further, a random-forest model trained on a combination of both datasets shows that cobalt-rich and iron-rich materials have the highest Curie temperatures for all binary and ternary compounds. An analysis of the model reveals systematic error that causes the model to over-predict low-$T$ C materials and under-predict high-$T$ C materials. For exhaustive searches to find new high-$T$ C materials, analysis of the learning rate suggests either that much more data is needed or that more efficient descriptors are necessary.

36 MATERIALS SCIENCE↗

Grad–Shafranov equilibria via data-free physics informed neural networks

A large number of magnetohydrodynamic (MHD) equilibrium calculations are often required for uncertainty quantification, optimization, and real-time diagnostic information, making MHD equilibrium codes vital to the field of plasma physics. In this paper, we explore a method for solving the Grad–Shafranov equation by using physics-informed neural networks (PINNs). For PINNs, we optimize neural networks by directly minimizing the residual of the partial differential equation as a loss function. We show that PINNs can accurately and effectively solve the Grad–Shafranov equation with several different boundary conditions, making it more flexible than traditional solvers. This method is flexible as it does not require any mesh and basis choice, thereby streamlining the computational process. We also explore the parameter space by varying the size of the model, the learning rate, and boundary conditions to map various tradeoffs such as between reconstruction error and computational speed. Additionally, we introduce a parameterized PINN framework, expanding the input space to include variables such as pressure, aspect ratio, elongation, and triangularity in order to handle a broader range of plasma scenarios within a single network. Parameterized PINNs could be used in future work to solve inverse problems such as shape optimization.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Group-equivariant autoencoder for identifying spontaneously broken symmetries

We introduce the group-equivariant autoencoder (GE autoencoder), a deep neural network (DNN) method that locates phase boundaries by determining which symmetries of the Hamiltonian have spontaneously broken at each temperature. We use group theory to deduce which symmetries of the system remain intact in all phases, and then use this information to constrain the parameters of the GE autoencoder such that the encoder learns an order parameter invariant to these “never-broken” symmetries. This procedure produces a dramatic reduction in the number of free parameters such that the GE-autoencoder size is independent of the system size. We include symmetry regularization terms in the loss function of the GE autoencoder so that the learned order parameter is also equivariant to the remaining symmetries of the system. By examining the group representation by which the learned order parameter transforms, we are then able to extract information about the associated spontaneous symmetry breaking. We test the GE autoencoder on the 2D classical ferromagnetic and antiferromagnetic Ising models, finding that the GE autoencoder (1) accurately determines which symmetries have spontaneously broken at each temperature; (2) estimates the critical temperature in the thermodynamic limit with greater accuracy, robustness, and time efficiency than a symmetry-agnostic baseline autoencoder; and (3) detects the presence of an external symmetry-breaking magnetic field with greater sensitivity than the baseline method. Lastly, we describe various key implementation details, including a quadratic-programming-based method for extracting the critical temperature estimate from trained autoencoders and calculations of the DNN initialization and learning rate settings required for fair model comparisons.

42 ENGINEERING↗

Optimal Balance of Privacy and Utility with Differential Privacy Deep Learning Frameworks

As the number of online services has increased, the amount of sensitive data being recorded is rising. Simultaneously, the decision-making process has improved by using the vast amounts of data, where machine learning has transformed entire industries. This paper addresses the development of optimal private deep neural networks and discusses the challenges associated with this task. We focus on differential privacy implementations and finding the optimal balance between accuracy and privacy, benefits and limitations of existing libraries, and challenges of applying private machine learning models in practical applications. Our analysis shows that learning rate, and privacy budget are the key factors that impact the results, and we discuss options for these settings.

Kotevska, Olivera↗

Multiobjective Hyperparameter Optimization for Deep Learning Interatomic Potential Training Using NSGA-II

Deep neural network (DNN) potentials are an emerging tool for simulation of dynamical atomistic systems, with the promise of quantum mechanical accuracy at speedups of 10000$\times$. As with other DNN methods, hyperparameters used during training can make a substantial difference in model accuracy, and optimal settings vary with dataset. To enable rapid tuning of hyperparameters for DNN potential training, we developed a scalable multiobjective optimization evolutionary algorithm for supercomputers and tested it on the Summit system at the Oak Ridge Leadership Computing Facility (OLCF). The multiobjective approach is required due to the coupling of two learned values defining the potential: the energy and force. Using a large-scale implementation of the NSGA-II algorithm adapted for training DNN potentials, we discovered several optimal multiobjective combinations, including best choices of activation functions, learning rate scaling scheme, and pairing of the two radial cutoffs used in the three dimensional descriptor function.

Coletti, Mark↗

pnnl/brain_ohsu

Light sheet microscopy has made possible the 3D imaging of both fixed and live biological tissue, with samples as large as the entire mouse brain. We fine-tuned an existing model, TrailMap, using expert labeled data from axonal structures in neocortex. Without changing the network architecture, we implemented nnU-Net framework modifications in data augmentation, data foreground sampling, window learning rate, and the inference overlap method. The resulting model from these combined approaches yielded an improved F1 score

Oostrom, Marjolein↗

Toadstool Deep Learning Framework

SAND2025-11741O Toadstool is a deep learning framework and support library that provides PyTorch boiler plate training and testing loops. This enables the user to remember parts and customize a callback interface. Toadstool also provides useful callbacks and other methods for deep learning experimentation. The framework also implements publicly available temperature and calibration methods, model initialization methods, learning rate schedulers, and model evaluation methods. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Heidbrink, Scott↗

Improving ADAM through an implicit-explicit (IMEX) time-stepping approach

The ADAM optimizer, often used in machine learning for neural network training, corresponds to an underlying ordinary differential equation (ODE) in the limit of very small learning rates. Here, this work shows that the classical ADAM algorithm is a first-order implicit-explicit (IMEX) Euler discretization of the underlying ODE. Employing the time discretization point of view, we propose new extensions of the ADAM scheme obtained by using higher-order IMEX methods to solve the ODE. Based on this approach, we derive a new optimization algorithm for neural network training that performs better than classical ADAM on several regression and classification problems.

97 MATHEMATICS AND COMPUTING↗

An Economics-by-Design Approach Applied to a Heat Pipe Microreactor Concept

Microreactors present a potential paradigm shift in the nuclear industry. Emphasis thus far has been on large-scale multi-billion-dollar projects that cater solely to grid electricity market. These projects can be challenging to finance and execute. On the other hand, microreactors are intended to target a wide variety of smaller niche markets and are expected to be factory-fabricated and more readily deployable. While diseconomies of scale for microreactors may tend to raise their costs per energy output (MWh) relative to large nuclear plants, offsetting gains can be expected from standardization, simplification, passive safety, lower radionuclide inventories, factory fabrication, fast installation, and low financing costs. To adequately assess these contributions, designers should have a different perspective on cost drivers than for large nuclear plants and can utilize novel approaches for systematic cost reduction. To account for these important aspects of microreactors, this report proposes an economics-by-design approach that places economic considerations at the center of the design process. The methodology builds on existing frameworks such as design-to-cost and value engineering, expanding them to new markets (beyond the grid), new attributes (beyond costs alone), and introducing the approach at earlier points in the design cycle. Design parameters and technical specifications are systematically evaluated until costs meet market entry points, while also providing the high-priority performance attributes of the particular use case. Determining first-order estimates for different components early in the process enables designers to focus R&D efforts on the biggest overall cost contributors and components with the most cost uncertainty. The analysis is always guided by market needs and threshold prices. In addition to microreactors, the approach is expected to be useful for other classes of nuclear reactors as well. The analysis was applied to a concept found in the open literature (the Design A heat-pipe reactor). A comprehensive bottom-up estimate was generated by leveraging a new microreactor-specific code of accounts and a range of cost equations. The initial estimate for levelized cost of electricity (LCOE) unsurprisingly exceeded market ranges since the use case had prioritized technological readiness over economic considerations in design choices. An alternate concept was then proposed, with various assumptions/targets made to reduce the largest cost contributors. Changes in the neutron spectrum, the power output, and building structures were found to make even the first-of-a-kind of this modified concept competitive with diesel generation in some remote communities. Learning rate (LR) assumptions indicated cost reductions achieved from sequential unit deployments could expand the range of competitiveness to include additional markets as deployments proceed.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Diagnosing and overcoming recombination and resistive losses in non-silicon solar cells using a silicon-inspired characterization platform

This project aimed to generate the characterization tools needed for accurate and systematic loss analysis in non-silicon photovoltaic solar cell technologies and, through the use of these tools and analysis techniques, contribute to the development of a novel class of hetero-contacts to II-VI absorbers, with the final goal of demonstrating record-breaking CdSeTe devices. CdSeTe solar cells provide a prime example of the potential impact of the techniques we proposed to develop and implement: record poly-CdSeTe cells have bandgap-voltage deficits (W oc ) of approximately 550 mV, as compared with below 400 mV for all other mature PV technologies. Similarly, these record CdSeTe devices have FFs below 80%, when other mature cells are near or above 85%. Frustratingly, a systematic identification of the origin of these sub-par performances—for example recombination or resistive losses—has been lacking, thus slowing down the development of these technologies. Similarly, it is often asserted that CdSeTe cells need a better back (hole) contact. Although most believe this is true, no one knew—at the start of this project—how high the V oc and FF could be for a given cell if it had a perfect back contact. Such characterization techniques and loss analysis methods exist and are routinely performed on c-Si solar cells (e.g. injection-dependent lifetime, Suns-V oc , transfer length method, etc). Over the years, they have been instrumental in the development of silicon devices that operate at 91% of their theoretical (Auger) limit. Lifetime testing, and the associated reconstruction of the implied-J-V curve, can moreover be performed at every cell-processing step, thus allowing a direct peek into the impact of that step on cell performance. Therefore, adapting these techniques and tools to non-Si devices would greatly improve their learning rate. In this project, we developed a Suns-ERE technique—the equipment, methodology, and know-how—to measure the implied-J-V curve, the pseudo-J-V curve, and the actual J-V curve of a thin-film solar cell, allowing an accurate assessment of the quality of the bulk material and its surface passivation, the selectivity of the contact, and its resistivity. We used this technique to show that the absorber of present CdSeTe solar cells is capable of achieving 1 V Voc,, that passivation layers exist (e.g., Al2O3) that can support such high voltages, and that the barrier is identifying contact layers that are both passivating and carrier-selective. The characterization platform created in this project and the understanding generated using it will accelerate the progress of non-silicon PV technologies. In particular, the project will contribute to CdSeTe solar cells with Voc > 1 V and cell efficiency > 24%. Such cells provide a pathway to module-level efficiencies >23%. As CdSeTe presently competes with silicon on module cost (in $\$ $/W) and yet has significantly more room for efficiency gains, the potential for LCOE reduction is particularly large. For example, CdTe modules with an efficiency of 21% would allow an LCOE below $\$ $0.04kWh -1 in average US climates.

14 SOLAR ENERGY↗

Potential Fuel Cycle Cost Reductions of Once Through HALEU Reactors

This report looks to identify potential cost reduction opportunities for SFR and HTGR reactors using HALEU fuel. This analysis focuses on how learning rates and experience from other industries could translate to future HALEU fueled reactor fuel cycles. Various fuel loading and residence scenarios are evaluated to estimate areas of potential cost savings. Cost savings from location optimization is explored for both fresh and spent nuclear fuel. Additional analysis was completed to estimate cost savings through improved labor productivity.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Data for The utility of transfer learning to improve the performance of deep learning in axon segmentation

The utility of transfer learning to improve the performance of deep learning in axon segmentation Data Data: All the input and labeled volumes tf-logs: Tensorflow logs, view with command "tensorboard --logdir [name of folder]" Model Weights: model_weights: the argument list under variable combo indicate 1) no oversampling, 2) no rotation, 3) no learn scheduler, and 4) flipping on all three dimensions, and the additional values indicate 5) elastic deformation percentage, 6) rotate deformation percentage, 7) layer setting , 8) learning rate, and 9) training/validation/test data division suffix (leave '' if not using suffix). Results: Output from inference segment_total_results_validation_final: All validation results and calculations segment_total_results: All test results and calculations Authors The modified code was created for a paper by: Marjolein Oostrom, Michael A. Muniak, Rogene Eichler West, Sarah Akers, Paritosh Pande, Moses Obiri, Wei Wang, Kasey Bowyer, Zhuhao Wu, Lisa Bramer, Tianyi Mao, Bobbie Jo Webb-Robertson The work is adapted from Github TrailMap, which was created by Albert Pun and Drew Friedmann Acknowledgments MO, RMEW, SA, MO, LB, BJWR were supported by the Laboratory Directed Research and Development at Pacific Northwest National Laboratory (PNNL), a Department of Energy facility operated by Battelle under contract DE-AC05-76RLO01830. WW, KB, and ZW were supported in part by a NIH/BRAIN Initiative Grant RF1MH128969. MAM and TM were supported by two NIH/BRAIN Initiative Grants R01NS104944, RF1MH120119 and NIH R01NS081071. This research is affiliated with the Pacific northwest bioMedical Innovation Co-laboratory (PMedIC) collaboration between OHSU and PNNL.

Oostrom, Marjolein T↗