Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “learning rate”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Group-equivariant autoencoder for identifying spontaneously broken symmetries

We introduce the group-equivariant autoencoder (GE autoencoder), a deep neural network (DNN) method that locates phase boundaries by determining which symmetries of the Hamiltonian have spontaneously broken at each temperature. We use group theory to deduce which symmetries of the system remain intact in all phases, and then use this information to constrain the parameters of the GE autoencoder such that the encoder learns an order parameter invariant to these “never-broken” symmetries. This procedure produces a dramatic reduction in the number of free parameters such that the GE-autoencoder size is independent of the system size. We include symmetry regularization terms in the loss function of the GE autoencoder so that the learned order parameter is also equivariant to the remaining symmetries of the system. By examining the group representation by which the learned order parameter transforms, we are then able to extract information about the associated spontaneous symmetry breaking. We test the GE autoencoder on the 2D classical ferromagnetic and antiferromagnetic Ising models, finding that the GE autoencoder (1) accurately determines which symmetries have spontaneously broken at each temperature; (2) estimates the critical temperature in the thermodynamic limit with greater accuracy, robustness, and time efficiency than a symmetry-agnostic baseline autoencoder; and (3) detects the presence of an external symmetry-breaking magnetic field with greater sensitivity than the baseline method. Lastly, we describe various key implementation details, including a quadratic-programming-based method for extracting the critical temperature estimate from trained autoencoders and calculations of the DNN initialization and learning rate settings required for fair model comparisons.

42 ENGINEERING↗

Optimal Balance of Privacy and Utility with Differential Privacy Deep Learning Frameworks

As the number of online services has increased, the amount of sensitive data being recorded is rising. Simultaneously, the decision-making process has improved by using the vast amounts of data, where machine learning has transformed entire industries. This paper addresses the development of optimal private deep neural networks and discusses the challenges associated with this task. We focus on differential privacy implementations and finding the optimal balance between accuracy and privacy, benefits and limitations of existing libraries, and challenges of applying private machine learning models in practical applications. Our analysis shows that learning rate, and privacy budget are the key factors that impact the results, and we discuss options for these settings.

Kotevska, Olivera↗

Multiobjective Hyperparameter Optimization for Deep Learning Interatomic Potential Training Using NSGA-II

Deep neural network (DNN) potentials are an emerging tool for simulation of dynamical atomistic systems, with the promise of quantum mechanical accuracy at speedups of 10000$\times$. As with other DNN methods, hyperparameters used during training can make a substantial difference in model accuracy, and optimal settings vary with dataset. To enable rapid tuning of hyperparameters for DNN potential training, we developed a scalable multiobjective optimization evolutionary algorithm for supercomputers and tested it on the Summit system at the Oak Ridge Leadership Computing Facility (OLCF). The multiobjective approach is required due to the coupling of two learned values defining the potential: the energy and force. Using a large-scale implementation of the NSGA-II algorithm adapted for training DNN potentials, we discovered several optimal multiobjective combinations, including best choices of activation functions, learning rate scaling scheme, and pairing of the two radial cutoffs used in the three dimensional descriptor function.

Coletti, Mark↗

pnnl/brain_ohsu

Light sheet microscopy has made possible the 3D imaging of both fixed and live biological tissue, with samples as large as the entire mouse brain. We fine-tuned an existing model, TrailMap, using expert labeled data from axonal structures in neocortex. Without changing the network architecture, we implemented nnU-Net framework modifications in data augmentation, data foreground sampling, window learning rate, and the inference overlap method. The resulting model from these combined approaches yielded an improved F1 score

Oostrom, Marjolein↗

Toadstool Deep Learning Framework

SAND2025-11741O Toadstool is a deep learning framework and support library that provides PyTorch boiler plate training and testing loops. This enables the user to remember parts and customize a callback interface. Toadstool also provides useful callbacks and other methods for deep learning experimentation. The framework also implements publicly available temperature and calibration methods, model initialization methods, learning rate schedulers, and model evaluation methods. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Heidbrink, Scott↗

Improving ADAM through an implicit-explicit (IMEX) time-stepping approach

The ADAM optimizer, often used in machine learning for neural network training, corresponds to an underlying ordinary differential equation (ODE) in the limit of very small learning rates. Here, this work shows that the classical ADAM algorithm is a first-order implicit-explicit (IMEX) Euler discretization of the underlying ODE. Employing the time discretization point of view, we propose new extensions of the ADAM scheme obtained by using higher-order IMEX methods to solve the ODE. Based on this approach, we derive a new optimization algorithm for neural network training that performs better than classical ADAM on several regression and classification problems.

97 MATHEMATICS AND COMPUTING↗

An Economics-by-Design Approach Applied to a Heat Pipe Microreactor Concept

Microreactors present a potential paradigm shift in the nuclear industry. Emphasis thus far has been on large-scale multi-billion-dollar projects that cater solely to grid electricity market. These projects can be challenging to finance and execute. On the other hand, microreactors are intended to target a wide variety of smaller niche markets and are expected to be factory-fabricated and more readily deployable. While diseconomies of scale for microreactors may tend to raise their costs per energy output (MWh) relative to large nuclear plants, offsetting gains can be expected from standardization, simplification, passive safety, lower radionuclide inventories, factory fabrication, fast installation, and low financing costs. To adequately assess these contributions, designers should have a different perspective on cost drivers than for large nuclear plants and can utilize novel approaches for systematic cost reduction. To account for these important aspects of microreactors, this report proposes an economics-by-design approach that places economic considerations at the center of the design process. The methodology builds on existing frameworks such as design-to-cost and value engineering, expanding them to new markets (beyond the grid), new attributes (beyond costs alone), and introducing the approach at earlier points in the design cycle. Design parameters and technical specifications are systematically evaluated until costs meet market entry points, while also providing the high-priority performance attributes of the particular use case. Determining first-order estimates for different components early in the process enables designers to focus R&D efforts on the biggest overall cost contributors and components with the most cost uncertainty. The analysis is always guided by market needs and threshold prices. In addition to microreactors, the approach is expected to be useful for other classes of nuclear reactors as well. The analysis was applied to a concept found in the open literature (the Design A heat-pipe reactor). A comprehensive bottom-up estimate was generated by leveraging a new microreactor-specific code of accounts and a range of cost equations. The initial estimate for levelized cost of electricity (LCOE) unsurprisingly exceeded market ranges since the use case had prioritized technological readiness over economic considerations in design choices. An alternate concept was then proposed, with various assumptions/targets made to reduce the largest cost contributors. Changes in the neutron spectrum, the power output, and building structures were found to make even the first-of-a-kind of this modified concept competitive with diesel generation in some remote communities. Learning rate (LR) assumptions indicated cost reductions achieved from sequential unit deployments could expand the range of competitiveness to include additional markets as deployments proceed.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Diagnosing and overcoming recombination and resistive losses in non-silicon solar cells using a silicon-inspired characterization platform

This project aimed to generate the characterization tools needed for accurate and systematic loss analysis in non-silicon photovoltaic solar cell technologies and, through the use of these tools and analysis techniques, contribute to the development of a novel class of hetero-contacts to II-VI absorbers, with the final goal of demonstrating record-breaking CdSeTe devices. CdSeTe solar cells provide a prime example of the potential impact of the techniques we proposed to develop and implement: record poly-CdSeTe cells have bandgap-voltage deficits (W oc ) of approximately 550 mV, as compared with below 400 mV for all other mature PV technologies. Similarly, these record CdSeTe devices have FFs below 80%, when other mature cells are near or above 85%. Frustratingly, a systematic identification of the origin of these sub-par performances—for example recombination or resistive losses—has been lacking, thus slowing down the development of these technologies. Similarly, it is often asserted that CdSeTe cells need a better back (hole) contact. Although most believe this is true, no one knew—at the start of this project—how high the V oc and FF could be for a given cell if it had a perfect back contact. Such characterization techniques and loss analysis methods exist and are routinely performed on c-Si solar cells (e.g. injection-dependent lifetime, Suns-V oc , transfer length method, etc). Over the years, they have been instrumental in the development of silicon devices that operate at 91% of their theoretical (Auger) limit. Lifetime testing, and the associated reconstruction of the implied-J-V curve, can moreover be performed at every cell-processing step, thus allowing a direct peek into the impact of that step on cell performance. Therefore, adapting these techniques and tools to non-Si devices would greatly improve their learning rate. In this project, we developed a Suns-ERE technique—the equipment, methodology, and know-how—to measure the implied-J-V curve, the pseudo-J-V curve, and the actual J-V curve of a thin-film solar cell, allowing an accurate assessment of the quality of the bulk material and its surface passivation, the selectivity of the contact, and its resistivity. We used this technique to show that the absorber of present CdSeTe solar cells is capable of achieving 1 V Voc,, that passivation layers exist (e.g., Al2O3) that can support such high voltages, and that the barrier is identifying contact layers that are both passivating and carrier-selective. The characterization platform created in this project and the understanding generated using it will accelerate the progress of non-silicon PV technologies. In particular, the project will contribute to CdSeTe solar cells with Voc > 1 V and cell efficiency > 24%. Such cells provide a pathway to module-level efficiencies >23%. As CdSeTe presently competes with silicon on module cost (in $\$ $/W) and yet has significantly more room for efficiency gains, the potential for LCOE reduction is particularly large. For example, CdTe modules with an efficiency of 21% would allow an LCOE below $\$ $0.04kWh -1 in average US climates.

14 SOLAR ENERGY↗

Potential Fuel Cycle Cost Reductions of Once Through HALEU Reactors

This report looks to identify potential cost reduction opportunities for SFR and HTGR reactors using HALEU fuel. This analysis focuses on how learning rates and experience from other industries could translate to future HALEU fueled reactor fuel cycles. Various fuel loading and residence scenarios are evaluated to estimate areas of potential cost savings. Cost savings from location optimization is explored for both fresh and spent nuclear fuel. Additional analysis was completed to estimate cost savings through improved labor productivity.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Data for The utility of transfer learning to improve the performance of deep learning in axon segmentation

The utility of transfer learning to improve the performance of deep learning in axon segmentation Data Data: All the input and labeled volumes tf-logs: Tensorflow logs, view with command "tensorboard --logdir [name of folder]" Model Weights: model_weights: the argument list under variable combo indicate 1) no oversampling, 2) no rotation, 3) no learn scheduler, and 4) flipping on all three dimensions, and the additional values indicate 5) elastic deformation percentage, 6) rotate deformation percentage, 7) layer setting , 8) learning rate, and 9) training/validation/test data division suffix (leave '' if not using suffix). Results: Output from inference segment_total_results_validation_final: All validation results and calculations segment_total_results: All test results and calculations Authors The modified code was created for a paper by: Marjolein Oostrom, Michael A. Muniak, Rogene Eichler West, Sarah Akers, Paritosh Pande, Moses Obiri, Wei Wang, Kasey Bowyer, Zhuhao Wu, Lisa Bramer, Tianyi Mao, Bobbie Jo Webb-Robertson The work is adapted from Github TrailMap, which was created by Albert Pun and Drew Friedmann Acknowledgments MO, RMEW, SA, MO, LB, BJWR were supported by the Laboratory Directed Research and Development at Pacific Northwest National Laboratory (PNNL), a Department of Energy facility operated by Battelle under contract DE-AC05-76RLO01830. WW, KB, and ZW were supported in part by a NIH/BRAIN Initiative Grant RF1MH128969. MAM and TM were supported by two NIH/BRAIN Initiative Grants R01NS104944, RF1MH120119 and NIH R01NS081071. This research is affiliated with the Pacific northwest bioMedical Innovation Co-laboratory (PMedIC) collaboration between OHSU and PNNL.

Oostrom, Marjolein T↗

An Accuracy-Maximization Approach for Claims Classifiers in Document Content Analytics for Cybersecurity

This paper presents our research approach and findings towards maximizing the accuracy of our classifier of feature claims for cybersecurity literature analytics, and introduces the resulting model ClaimsBERT. Its architecture, after extensive evaluations of different approaches, introduces a feature map concatenated with a Bidirectional Encoder Representation from Transformers (BERT) model. We discuss deployment of this new concept and the research insights that resulted in the selection of Convolution Neural Networks for its feature mapping aspects. We also present our results showing ClaimsBERT to outperform all other evaluated approaches. This new claims classifier represents an essential processing stage within our vetting framework aiming to improve the cybersecurity of industrial control systems (ICS). Furthermore, in order to maximize the accuracy of our new ClaimsBERT classifier, we propose an approach for optimal architecture selection and determination of optimized hyperparameters, in particular the best learning rate, number of convolutions, filter sizes, activation function, the number of dense layers, as well as the number of neurons and the drop-out rate for each layer. Fine-tuning these hyperparameters within our model led to an increase in classification accuracy from 76% obtained with BertForSequenceClassification’s original model to a 97% accuracy obtained with ClaimsBERT.

Ameri, Kimia (ORCID:0000000328791871)↗

Space shuttle solid rocket booster cost-per-flight analysis technique

A cost per flight computer model is described which considers: traffic model, component attrition, hardware useful life, turnaround time for refurbishment, manufacturing rates, learning curves on the time to perform tasks, cost improvement curves on quantity hardware buys, inflation, spares philosophy, long lead, hardware funding requirements, and other logistics and scheduling constraints. Additional uses of the model include assessing the cost per flight impact of changing major space shuttle program parameters and searching for opportunities to make cost effective management decisions.

Forney, J. A.↗

Satellite servicing economic study

Previous studies have shown that satellite servicing is cost effective; however, all of these studies were of different formats, dollar year, learning rates, availability, etc. Therefore, it was difficult to correlate any useful trends from these studies. The reviewed study was initiated to correlate the economic data into a common data base, using a common set of assumptions. A selected set of existed funded programs was then analyzed to provide an independent analysis of the servicing options and potential economic benefits.

Source record↗

Satellite servicing economic study

Previous studies have shown that satellite servicing is cost effective; however, all of these studies were of different formats, dollar year, learning rates, availability, etc. Threfore, it was difficult to correlate any useful trends from these studies. The reviewed study was initiated to correlate the economic data into a common data base, using a common set of assumptions. A selected set of existed funded programs was then analyzed to provide an independent analysis of the servicing options and potential economic benefits.

Source record↗

Learning reliable manipulation strategies without initial physical models

A description is given of a robot, possessing limited sensory and effectory capabilities but no initial model of the effects of its actions on the world, that acquires such a model through exploration, practice, and observation. By acquiring an increasingly correct model of its actions, it generates increasingly successful plans to achieve its goals. In an apparently nondeterministic world, achieving reliability requires the identification of reliable actions and a preference for using such actions. Furthermore, by selecting its training actions carefully, the robot can significantly improve its learning rate.

Christiansen, Alan D.↗

Learning characteristics of a space-time neural network as a tether skiprope observer

The Software Technology Laboratory at the Johnson Space Center is testing a Space Time Neural Network (STNN) for observing tether oscillations present during retrieval of a tethered satellite. Proper identification of tether oscillations, known as 'skiprope' motion, is vital to safe retrieval of the tethered satellite. Our studies indicate that STNN has certain learning characteristics that must be understood properly to utilize this type of neural network for the tethered satellite problem. We present our findings on the learning characteristics including a learning rate versus momentum performance table.

Lea, Robert N.↗

Two generalizations of Kohonen clustering

The relationship between the sequential hard c-means (SHCM), learning vector quantization (LVQ), and fuzzy c-means (FCM) clustering algorithms is discussed. LVQ and SHCM suffer from several major problems. For example, they depend heavily on initialization. If the initial values of the cluster centers are outside the convex hull of the input data, such algorithms, even if they terminate, may not produce meaningful results in terms of prototypes for cluster representation. This is due in part to the fact that they update only the winning prototype for every input vector. The impact and interaction of these two families with Kohonen's self-organizing feature mapping (SOFM), which is not a clustering method, but which often leads ideas to clustering algorithms is discussed. Then two generalizations of LVQ that are explicitly designed as clustering algorithms are presented; these algorithms are referred to as generalized LVQ = GLVQ; and fuzzy LVQ = FLVQ. Learning rules are derived to optimize an objective function whose goal is to produce 'good clusters'. GLVQ/FLVQ (may) update every node in the clustering net for each input vector. Neither GLVQ nor FLVQ depends upon a choice for the update neighborhood or learning rate distribution - these are taken care of automatically. Segmentation of a gray tone image is used as a typical application of these algorithms to illustrate the performance of GLVQ/FLVQ.

Bezdek, James C.↗

Image segmentation using fuzzy LVQ clustering networks

In this note we formulate image segmentation as a clustering problem. Feature vectors extracted from a raw image are clustered into subregions, thereby segmenting the image. A fuzzy generalization of a Kohonen learning vector quantization (LVQ) which integrates the Fuzzy c-Means (FCM) model with the learning rate and updating strategies of the LVQ is used for this task. This network, which segments images in an unsupervised manner, is thus related to the FCM optimization problem. Numerical examples on photographic and magnetic resonance images are given to illustrate this approach to image segmentation.

Tsao, Eric Chen-Kuo↗