Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “learning rate”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Automatic learning rate adjustment for self-supervising autonomous robot control

Described is an application in which an Artificial Neural Network (ANN) controls the positioning of a robot arm with five degrees of freedom by using visual feedback provided by two cameras. This application and the specific ANN model, local liner maps, are based on the work of Ritter, Martinetz, and Schulten. We extended their approach by generating a filtered, average positioning error from the continuous camera feedback and by coupling the learning rate to this error. When the network learns to position the arm, the positioning error decreases and so does the learning rate until the system stabilizes at a minimum error and learning rate. This abolishes the need for a predetermined cooling schedule. The automatic cooling procedure results in a closed loop control with no distinction between a learning phase and a production phase. If the positioning error suddenly starts to increase due to an internal failure such as a broken joint, or an environmental change such as a camera moving, the learning rate increases accordingly. Thus, learning is automatically activated and the network adapts to the new condition after which the error decreases again and learning is 'shut off'. The automatic cooling is therefore a prerequisite for the autonomy and the fault tolerance of the system.

Arras, Michael K.↗

Open Architecture for Cost Savings in Advanced Nuclear Reactors

Recently, nuclear power plant build projects in the West have run over budget due to high capital costs and schedule overruns. Compared to other sources of energy, nuclear power plants have higher capital costs. Reactors are often different at every site, resulting in a lack of standardization. Nuclear is expected to compete with other low carbon sources of energy which have lower capital costs making it essential for nuclear to develop ways of reducing costs. Strategies such as standardization, learning rates, modularization, and schedule reduction in advanced reactors can reduce nuclear costs by about 40%. Standardization as a way of cutting capital costs has been explored even in large nuclear power plants. Standardization of certain plant components can result in lower component and installation costs and higher learning from experience. Standardization can be achieved by adopting a criterion of key performance indicators and general design principles for a specific system or component such as the balance of plant. Modularization allows the construction of certain components of SMRs in a factory, which saves time, increases productivity, and encourages higher learning rates. Production learning decreases the time and the cost related to an activity. The potential for modularized components of advanced reactors to be manufactured in factories makes it conducive to achieving higher learning rates. Developing large-capacity nuclear programs through sequential builds cultivates a higher learning rate, which in effect may reduce schedule overruns. Open architecture has been identified as a way to drive standardization among advanced reactor designs and result in cost savings. Open architecture (OA) is defined as a design enabling a diverse supply chain by defining and publishing requirements of systems or equipment in functional and/or interface terms, utilizing technical standards in widespread use. Currently, the nuclear industry’s approach is to use closed architecture, making most designs proprietary. However, collaboration between various advanced reactor vendors and suppliers utilizing the concept of open architecture can result in modular and standardized architecture of subsystems or subcomponents of a nuclear power plant. Completely standardizing nuclear power plants may be impossible, however, certain common subsystems amongst the various reactor designs could be standardized and/or access a wider supply chain and leverage existing learning from other sectors. Open architecture will save time and allocate resources to the parts of the plants that have the most unique features. A key advantage of open architecture is its ability to improve production learning across advanced reactors (AR) types in the industry, by providing and utilizing the same kind of component. Sodium fast reactor (SFR), High Temperature Gas Reactor (HTGR) and Molten Salt Reactor (MSR) are the advanced reactors considered for this project. This paper aims to determine the cost savings in advanced reactor programs due to open architecture learning rate. This work is an extension of work done on light water reactor small modular reactors; the cost methodology was utilized to investigate the impact of open architecture on advanced reactors with a particular focus on sodium fast reactors. The cost data on sodium fast reactors used in the model presented the most adequate information required for the analysis.

Advanced Nuclear Reactors↗

Method of Real-Time Principal-Component Analysis

Dominant-element-based gradient descent and dynamic initial learning rate (DOGEDYN) is a method of sequential principal-component analysis (PCA) that is well suited for such applications as data compression and extraction of features from sets of data. In comparison with a prior method of gradient-descent-based sequential PCA, this method offers a greater rate of learning convergence. Like the prior method, DOGEDYN can be implemented in software. However, the main advantage of DOGEDYN over the prior method lies in the facts that it requires less computation and can be implemented in simpler hardware. It should be possible to implement DOGEDYN in compact, low-power, very-large-scale integrated (VLSI) circuitry that could process data in real time.

Duong, Tuan↗

Meta-Analysis of Advanced Nuclear Reactor Cost Estimations

Supporting Data can be downloaded at: https://gain.inl.gov/content/uploads/4/2024/06/INL-RPT-24-77048-R1.xlsx Nuclear energy is a critical cornerstone of the current United States clean energy supply and may play a larger role in the future in support of a transition to a net-zero economy. The current fleet of nuclear reactors predominantly consists of large light-water reactors (LWRs), while many of the reactor designs under consideration are smaller and/or different technologies. Because these new designs have not yet been built, there is a high degree of uncertainty associated with their cost. This complicates energy-planning efforts because cost projections are not always standardized, consistent, and centralized in an easily accessible location. To help support energy planning in the US, this report provides advanced nuclear cost ranges using a transparent methodology along with other relevant information that can be used to help support decision making and energy planning. The purpose of this work was to conduct a methodical process for cost evaluation using only public information that was vetted with the end-goal to provide reference cost projections for nuclear energy. To provide a solid basis for these values, the approach and assumptions are explicitly laid out throughout the report allowing any user of the data to challenge or reconsider them. Because future US nuclear-reactor costs are still unknown due to little recent observed data, the report opted to compile a comprehensive list of bottom-up estimates and evaluate averages/trends within the data to identify reference ranges. This was deemed preferable to opining on the robustness or validity of one cost estimation versus another. To that end, the work evaluated thousands of lines of cost subaccounts from several bottom-up cost estimates. A wide variety of different reactor types captured in the data are of various sizes and technologies. Some of these reactors will be representative of advanced reactors under development while others will not. Thus, the results here are dependent on the data that are available and the accuracy of the estimates that are used. Each bottom-up estimate was reviewed to determine whether it was complete. Incomplete data sets were corrected to ensure an adequate basis of cross-comparison. The report is not without limitations and should be interpreted as an initial step to develop cost ranges for nuclear technology. Ultimately, future work can build upon the methodology with refined cost estimates to reduce uncertainty. US-based overnight capital cost (OCC) estimates were compiled from extensive data sets into ranges for both large and small reactor sizes for 2030. To project the cost declines over time, learning rates were sampled from literature sources. No SMRs were previously built; hence, learning rates based on bottom-up approaches (e.g., by quantifying the impact stemming from fabrication of different components, modular work, site construction, commissioning) were prioritized. For larger reactors, actual learning rates from deployments were used to project future costs (adjusted to account for standardization or lack thereof between designs). Other costs included are fixed and variable operations and maintenance costs. The final variables were capacity factors and ramp rates to support energy planning.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Real-Time Principal-Component Analysis

A recently written computer program implements dominant-element-based gradient descent and dynamic initial learning rate (DOGEDYN), which was described in Method of Real-Time Principal-Component Analysis (NPO-40034) NASA Tech Briefs, Vol. 29, No. 1 (January 2005), page 59. To recapitulate: DOGEDYN is a method of sequential principal-component analysis (PCA) suitable for such applications as data compression and extraction of features from sets of data. In DOGEDYN, input data are represented as a sequence of vectors acquired at sampling times. The learning algorithm in DOGEDYN involves sequential extraction of principal vectors by means of a gradient descent in which only the dominant element is used at each iteration. Each iteration includes updating of elements of a weight matrix by amounts proportional to a dynamic initial learning rate chosen to increase the rate of convergence by compensating for the energy lost through the previous extraction of principal components. In comparison with a prior method of gradient-descent-based sequential PCA, DOGEDYN involves less computation and offers a greater rate of learning convergence. The sequential DOGEDYN computations require less memory than would parallel computations for the same purpose. The DOGEDYN software can be executed on a personal computer.

Duong, Vu↗

Locomotion training of legged robots using hybrid machine learning techniques

In this study artificial neural networks and fuzzy logic are used to control the jumping behavior of a three-link uniped robot. The biped locomotion control problem is an increment of the uniped locomotion control. Study of legged locomotion dynamics indicates that a hierarchical controller is required to control the behavior of a legged robot. A structured control strategy is suggested which includes navigator, motion planner, biped coordinator and uniped controllers. A three-link uniped robot simulation is developed to be used as the plant. Neurocontrollers were trained both online and offline. In the case of on-line training, a reinforcement learning technique was used to train the neurocontroller to make the robot jump to a specified height. After several hundred iterations of training, the plant output achieved an accuracy of 7.4%. However, when jump distance and body angular momentum were also included in the control objectives, training time became impractically long. In the case of off-line training, a three-layered backpropagation (BP) network was first used with three inputs, three outputs and 15 to 40 hidden nodes. Pre-generated data were presented to the network with a learning rate as low as 0.003 in order to reach convergence. The low learning rate required for convergence resulted in a very slow training process which took weeks to learn 460 examples. After training, performance of the neurocontroller was rather poor. Consequently, the BP network was replaced by a Cerebeller Model Articulation Controller (CMAC) network. Subsequent experiments described in this document show that the CMAC network is more suitable to the solution of uniped locomotion control problems in terms of both learning efficiency and performance. A new approach is introduced in this report, viz., a self-organizing multiagent cerebeller model for fuzzy-neural control of uniped locomotion is suggested to improve training efficiency. This is currently being evaluated for a possible patent by NASA, Johnson Space Center. An alternative modular approach is also developed which uses separate controllers for each stage of the running stride. A self-organizing fuzzy-neural controller controls the height, distance and angular momentum of the stride. A CMAC-based controller controls the movement of the leg from the time the foot leaves the ground to the time of landing. Because the leg joints are controlled at each time step during flight, movement is smooth and obstacles can be avoided. Initial results indicate that this approach can yield fast, accurate results.

Simon, William E.↗

A Systematic Framework for Projecting the Future Cost of Offshore Wind Energy

Offshore wind costs are expected to decline rapidly in the short and medium term future as the industry grows and gains experience in manufacturing, installing, and operating commercial scale projects. Estimating the future costs of offshore wind energy is critical for evaluating the technology's economic performance, how it can fit into a broader clean energy economy, and how R&D investment can be allocated to advance the technology. We present a newly developed approach for forecasting these costs which focuses on an empirically-derived learning rate for capital costs and prescribed cost and performance improvements for operational costs and capacity factor. We establish baseline costs for a series of reference fixed-bottom and floating projects in 2021 and project cost trajectories to 2035, presenting both an average cost trajectory as well as describing the range of potential future costs associated with site-specific cost variations and uncertainty in the estimate of the learning rate. We also conduct sensitivity analyses showing the impact of different global deployments by 2035 and variations in the prescribed operational costs and capacity factors. The results show that fixed-bottom and floating offshore wind capital costs could decrease to around $\$$2,400/kW and $\$$3,300/kW by 2035, respectively, with ranges of $\$$2,100/kW - $\$$2,750/kW for fixed-bottom projects and $\$$2,850/kW - $\$$5,500/kW for floating projects. The levelized cost of energy of fixed-bottom and floating wind projects could decrease to $\$$53.1/MWh and $\$$63.9/MWh by 2035, with ranges of $\$$48.4/MWh - $\$$59.7/MWh for fixed-bottom projects and $\$$46.5/MWh - $\$$99.9/MWh for floating projects. By presenting the uncertainty associated with the forecast we provide a transparent description of the spectrum of potential cost trajectories for offshore wind.

17 WIND ENERGY↗

Active Learning with Irrelevant Examples

Active learning algorithms attempt to accelerate the learning process by requesting labels for the most informative items first. In real-world problems, however, there may exist unlabeled items that are irrelevant to the user's classification goals. Queries about these points slow down learning because they provide no information about the problem of interest. We have observed that when irrelevant items are present, active learning can perform worse than random selection, requiring more time (queries) to achieve the same level of accuracy. Therefore, we propose a novel approach, Relevance Bias, in which the active learner combines its default selection heuristic with the output of a simultaneously trained relevance classifier to favor items that are likely to be both informative and relevant. In our experiments on a real-world problem and two benchmark datasets, the Relevance Bias approach significantly improved the learning rate of three different active learning approaches.

machine learning↗

BUTTER - Empirical Deep Learning Dataset

The BUTTER Empirical Deep Learning Dataset represents an empirical study of the deep learning phenomena on dense fully connected networks, scanning across thirteen datasets, eight network shapes, fourteen depths, twenty-three network sizes (number of trainable parameters), four learning rates, six minibatch sizes, four levels of label noise, and fourteen levels of L1 and L2 regularization each. Multiple repetitions (typically 30, sometimes 10) of each combination of hyperparameters were preformed, and statistics including training and test loss (using a 80% / 20% shuffled train-test split) are recorded at the end of each training epoch. In total, this dataset covers 178 thousand distinct hyperparameter settings ("experiments"), 3.55 million individual training runs (an average of 20 repetitions of each experiments), and a total of 13.3 billion training epochs (three thousand epochs were covered by most runs). Accumulating this dataset consumed 5,448.4 CPU core-years, 17.8 GPU-years, and 111.2 node-years.

Array↗

Gravity Well Commercial Economics Assessment: Potential Revenue and Cost: Cooperative Research and Development (Final Report)

In the Gravity Well Revenue Study, we evaluate the potential revenue from energy storage using historical energy-only electricity prices, forward-looking projections of hourly electricity prices, and actual reported revenue. This analysis examines the impact of storage duration and round-trip efficiency, as well as the location of the storage, on storage revenue within the current and projected U.S. power system. We also investigated the impact of round-trip efficiency on storage revenue. We found that the relationship between storage revenue and round-trip efficiency is nonlinear. The value of improved round-trip efficiency declines as round-trip efficiency increases. In the Gravity Well Future Cost Study, we applied learning curves to predict the future cost trajectory of Gravity Wells (GrWs). Two types of analysis were implemented. The first was a bottom-up analysis that used historical learning rates for cost components, such as motors and gearboxes, and cost categories (e.g., engineering and design, etc.) to determine the learning-by-doing based single-factor learning curve. The single factor learning curve expresses the relationship between the cost of GrW and the number of units deployed (or the cumulative capacity). In the second analysis, we predicted future GrW costs via a top-down approach. This approach accounts for historical cost trends in other renewable energy and storage technologies, which have similarities with GrWs. Using a multifactor learning curve that accounts for both intrinsic (cumulative capacity) and extrinsic (the elasticity in the price of steel) factors, we estimated the future cost of GrWs.

25 ENERGY STORAGE↗

The Relationship Between Fidelity and Learning in Aviation Training and Assessment

Flight simulators can be designed to train pilots or assess their flight performance. Low-Fidelity simulators maximize the initial learning rate of novice pilots and minimize initial costs; whereas, expensive, high-fidelity simulators predict the realworld in-flight performance of expert pilots (Fink & Shriver, 1978 Hays & Singer 1989; Kinkade & Wheaton. 1972). Although intuitively appealing and intellectually convenient to generalize concepts of learning and assessment, what holds true for the role of fidelity in assessment may not always hold true for learning, and vice versa. To bring clarity to this issue, the author distinguishes the role of fidelity in learning from its role in assessment as a function of skill level by applying the hypothesis of Alessi (1988) and reviewing the Laughery, Ditzian, and Houtman (1982) study on simulator validity. Alessi hypothesized that there is it point beyond which one additional unit of flight-simulator fidelity results in a diminished rate of learning. The author of this current paper also suggests the existence of an optimal point beyond which one additional unit of flight-simulator fidelity results in a diminished rate of practical assessment of nonexpert pilot performance.

Noble, Cliff↗

Development of Advanced Verification and Validation Procedures and Tools for the Certification of Learning Systems in Aerospace Applications

Adaptive control technologies that incorporate learning algorithms have been proposed to enable automatic flight control and vehicle recovery, autonomous flight, and to maintain vehicle performance in the face of unknown, changing, or poorly defined operating environments. In order for adaptive control systems to be used in safety-critical aerospace applications, they must be proven to be highly safe and reliable. Rigorous methods for adaptive software verification and validation must be developed to ensure that control system software failures will not occur. Of central importance in this regard is the need to establish reliable methods that guarantee convergent learning, rapid convergence (learning) rate, and algorithm stability. This paper presents the major problems of adaptive control systems that use learning to improve performance. The paper then presents the major procedures and tools presently developed or currently being developed to enable the verification, validation, and ultimate certification of these adaptive control systems. These technologies include the application of automated program analysis methods, techniques to improve the learning process, analytical methods to verify stability, methods to automatically synthesize code, simulation and test methods, and tools to provide on-line software assurance.

Jacklin, Stephen↗

Fuzzy self-learning control for magnetic servo system

It is known that an effective control system is the key condition for successful implementation of high-performance magnetic servo systems. Major issues to design such control systems are nonlinearity; unmodeled dynamics, such as secondary effects for copper resistance, stray fields, and saturation; and that disturbance rejection for the load effect reacts directly on the servo system without transmission elements. One typical approach to design control systems under these conditions is a special type of nonlinear feedback called gain scheduling. It accommodates linear regulators whose parameters are changed as a function of operating conditions in a preprogrammed way. In this paper, an on-line learning fuzzy control strategy is proposed. To inherit the wealth of linear control design, the relations between linear feedback and fuzzy logic controllers have been established. The exercise of engineering axioms of linear control design is thus transformed into tuning of appropriate fuzzy parameters. Furthermore, fuzzy logic control brings the domain of candidate control laws from linear into nonlinear, and brings new prospects into design of the local controllers. On the other hand, a self-learning scheme is utilized to automatically tune the fuzzy rule base. It is based on network learning infrastructure; statistical approximation to assign credit; animal learning method to update the reinforcement map with a fast learning rate; and temporal difference predictive scheme to optimize the control laws. Different from supervised and statistical unsupervised learning schemes, the proposed method learns on-line from past experience and information from the process and forms a rule base of an FLC system from randomly assigned initial control rules.

Tarn, J. H.↗

Towards provably efficient quantum algorithms for large-scale machine-learning models

Large machine learning models are revolutionary technologies of artificial intelligence whose bottlenecks include huge computational expenses, power, and time used both in the pre-training and fine-tuning process. In this work, we show that fault-tolerant quantum computing could possibly provide provably efficient resolutions for generic (stochastic) gradient descent algorithms, scaling as $\mathcal{O}$(T 2 x polylog($n$)), where n is the size of the models and T is the number of iterations in the training, as long as the models are both sufficiently dissipative and sparse, with small learning rates. Based on earlier efficient quantum algorithms for dissipative differential equations, we find and prove that similar algorithms work for (stochastic) gradient descent, the primary algorithm for machine learning. In practice, we benchmark instances of large machine learning models from 7 million to 103 million parameters. We find that, in the context of sparse training, a quantum enhancement is possible at the early stage of learning after model pruning, motivating a sparse parameter download and re-upload scheme. Our work shows solidly that fault-tolerant quantum algorithms could potentially contribute to most state-of-the-art, large-scale machine-learning problems.

97 MATHEMATICS AND COMPUTING↗

Automated Pneumothorax Diagnosis using Deep Neural Networks

Thoracic ultrasound can provide information leading to rapid diagnosis of pneumothorax with improved accuracy over the standard physical examination and with higher sensitivity than anteroposterior chest radiography. However, the clinical We have Furthermore, remote environments, such as the battlefield or deep-space exploration, may lack expertise for diagnosing developed an automated image interpretation pipeline for the analysis of thoracic ultrasound data and the classification of pneumothorax events to provide decision support in such situations. Our pipeline consists of image preprocessing, data augmentation, and deep learning architectures for medical diagnosis. In this work, we demonstrate that robust, accurate interpretation of chest images and video can be achieved using deep neural networks. A number of novel image processing techniques were employed to achieve this result. Affine transformations were applied for data augmentation. Hyperparameters were optimized for learning rate, dropout regularization, batch size, and epoch iteration by a sequential model-based Bayesian approach. In addition, we utilized pretrained architecturesinterpretation of a patient medical image is highly operator dependent. certain pathologies., applying transfer learning and fine-tuning techniques to fully connected layers. Our pipeline yielded binary classification validation accuracies of 98.3% for M-mode images and 99.8% with B-mode video frames.

US Army collaboration↗

Surrogate Model Based Optimization for Finding Robust Deep Learning Model Architectures

Deep Learning (DL) models are increasingly used throughout the sciences. However, their performance and usefulness depend greatly on their architecture which is defined by hyperparameters such as the number of nodes, layers, the learning rate, etc. Tuning these hyperparameters is time-consuming because evaluating their performance requires a lengthy training step. Stochastic optimizers used in training lead to performance variability and potentially prediction reliability issues. In this talk, we will describe an automated optimization method based on surrogate models and active learning strategies for tuning DL model architectures. We take into account the prediction variability with the goal to identify architectures that make reliable and robust predictions. We demonstrate our developments on an application arising in particle physics.

deep learning↗