Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “empirical machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Diesel Passenger Vehicle Shares Influenced Covid-19 Changes in Urban Nitrogen Dioxide Pollution

Diesel-powered vehicles emit several times more nitrogen oxides than comparable gasoline-powered vehicles, leading to ambient nitrogen dioxide (NO2) pollution and adverse health impacts. The COVID-19 pandemic and ensuing changes in emissions provide a natural experiment to test whether NO2 reductions have been starker in regions of Europe with larger diesel passenger vehicle shares. Here we use a semi-empirical approach that combines in-situ NO2 observations from urban areas and an atmospheric composition model within a machine learning algorithm to estimate business-as-usual NO2 during the first wave of the COVID-19 pandemic in 2020. These estimates account for the moderating influences of meteorology, chemistry, and traffic. Comparing the observed NO2 concentrations against business-as-usual estimates indicates that diesel passenger vehicle shares played a major role in the magnitude of NO2 reductions. European cities with the five largest shares of diesel passenger vehicles experienced NO2 reductions ∼ 2.5 times larger than cities with the five smallest diesel shares. Extending our methods to a cohort of non-European cities reveals that NO2 reductions in these cities were generally smaller than reductions in European cities, which was expected given their small diesel shares. We identify potential factors such as the deterioration of engine controls associated with older diesel vehicles to explain spread in the relationship between cities’ shares of diesel vehicles and changes in NO2 during the pandemic. Our results provide a glimpse of potential NO2 reductions that could accompany future deliberate efforts to phase out or remove passenger vehicles from cities.

Nitrogen Dioxide↗

Hybrid Modeling for Complex Systems Health Management

The research work presents application of hybrid physics-informed machine learning to a representative electric powertrain for unmanned aerial vehicles. The model is composed of physics-derived and empirical equations, integrated with connected networks that are strategically placed within the model to substitute equations that are subject to large uncertainty. Polynomial fit driven by heuristics or empirical observations can be substituted by more flexible networks that can minimize the error between model predictions and observations without being restricted to a predefined functional form. This modeling strategy allows training of networks deep inside the model and unknown parameters in a single learning stage. The powertrain model consists of Li-ion batteries, electronic speed controller with pulse-width modulation, and brush-less DC motor with connected propeller. Results obtained from combination of laboratory and simulation tests are discussed in this work.

PINNS↗

Temperature Dependence of Band Gap Renormalization in High-T Sensor Materials via First-Principles and Experimental Corroboration

Understanding the temperature dependence of functional properties of high-T gas sensing materials is vital for their applications in combustion environments. The electron-phonon coupling that derives the electronic structure change with temperatures is a key property of interest as it affects other sensing responses. Herein, we assess the temperature dependence of band gap renormalization in metal oxides and perovskites by employing Allen-Heine-Cardona theory with first-principles simulations and corroborate with experimental observation. The calculated temperature-dependent band gap changes of these materials studied are in good agreement with in-house experimental data, proving that the theory can adequately predict renormalization on the band gap in the system of interest. The predicted and measured band gap variations are characterized using an analytical model, which can provide useful insights on the simulated zero-temperature band gaps. Based on the available data, a set of 53 metal oxides and perovskites were identified as potential high-T gas sensors. A machine learning model has been developed to predict the band-gap change by capturing the overall trend of the empirical parameters with respect to a reduced feature obtained by transforming the set of available physical features.

Park, Jongwoo↗

Deep learning of dynamically responsive chemical Hamiltonians with semiempirical quantum mechanics

Conventional machine-learning (ML) models in computational chemistry learn to directly predict molecular properties using quantum chemistry only for reference data. While these heuristic ML methods show quantum-level accuracy with speeds several orders of magnitude faster than traditional quantum chemistry methods, they suffer from poor extensibility and transferability; i.e., their accuracy degrades on large or new chemical systems. Incorporating quantum chemistry frameworks into the ML models directly solves this problem. Here we take the structure of semiempirical quantum mechanics (SEQM) methods to construct dynamically responsive Hamiltonians. SEQM methods use empirical parameters fitted to experimental properties to construct reduced-order Hamiltonians, facilitating much faster calculations than ab initio methods but with compromised accuracy. By replacing these static parameters with machine-learned dynamic values inferred from the local environment, we greatly improve the accuracy of the SEQM methods. Trained on molecular energies and atomic forces, these dynamically generated Hamiltonian parameters show a strong correlation with atomic hybridization and bonding. Trained with only about 60,000 small organic molecular conformers, the resulting model retains interpretability, extensibility, and transferability when testing on much larger chemical systems and predicting various molecular properties. Overall, this work demonstrates the virtues of incorporating physics-based descriptions with ML to develop models that are simultaneously accurate, transferable, and interpretable.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

An Applied Strategy for Using Empirical and Hybrid Models in Online Monitoring

The monitoring of plant equipment for failure prediction is one of the key contributors to operation and maintenance (O&M) costs for a nuclear power plant (NPP) because O&M monitoring depends on labor-intensive activities that are required to meet high equipment reliability standards. These activities rely primarily on humans for information gathering, condition diagnosis, and predictive analysis. Online monitoring aims to automate these activities by relying on sensors to replace human information gathering and machine learning to replace human analysis and decision making. To facilitate automated monitoring, a systematic strategy for anomaly detection is needed to optimally use the available sensor data, empirical models, and physics-supported models. This strategy is essential to provide credible reasoning on why and when an empirical (i.e., purely data-driven) versus hybrid (i.e., physics-supported) approach should be used and to determine the ideal mix of these two approaches for a defined anomaly detection scope. The extant methods usually adopt an ad hoc trial-and-error approach that, in addition to being time-consuming and costly, is also highly subjective; it is impacted by the background and the skill set of the personnel making the decisions. Thus, such an approach cannot guarantee an optimum outcome. This represents the motivation of the current research effort, which is focused on devising a scientifically supported strategy for the optimum selection of anomaly detection methods. This report presents a detailed assessment of the main anomaly detection techniques within the empirical or hybrid method streams. Empirical methods include pattern, statistical, and causal inference. Hybrid methods include the use of physics models to train and test data methods, reduce data dimensionality, reduce data-model complexity, augment data, and reduce empirical uncertainty; hybrid methods also include the use of data to tune physics models. The listed techniques within these two streams represent the vast majority of techniques performed for anomaly detection. Using the techniques as outcomes, a strategy was developed to enable a systematic decision-making process to lead to one of these techniques. The strategy is driven by key decision points related to data relevance, simple modeling feasibility, data inference, physics-modeling value, data dimensionality, physics knowledge, method of validation, performance, data availability and suitability for training and testing, cause-effect, entropy inference, and model fitting. Each of these decision points in the strategy is explained in detail in this report with examples, along with the scientific basis behind the decisions and outcomes in common and simplified terminology. The strategy is developed for use by any NPP staff with basic engineering or science knowledge. A user-friendly graphical state flow diagram was also developed as a visual presentation of the strategy. The strategy was tested and demonstrated through two pilot projects for the application of anomaly detection at an NPP. Each pilot had two use cases: an initial case in which certain decisions were made that resulted in one or more empirical techniques and a revised use case where one or more key decisions were modified resulting in using a set of hybrid methods.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

On the Convergence of Inexact Predictor-Corrector Methods for Linear Programming

Interior point methods (IPMs) are a common approach for solving linear programs (LPs) with strong theoretical guarantees and solid empirical performance. The time complexity of these methods is dominated by the cost of solving a linear system of equations at each iteration. In common applications of linear programming, particularly in machine learning and scientific computing, the size of this linear system can become prohibitively large, requiring the use of iterative solvers, which provide an approximate solution to the linear system. However, approximately solving the linear system at each iteration of an IPM invalidates the theoretical guarantees of common IPM analyses. To remedy this, we theoretically and empirically analyze (slightly modified) predictor-corrector IPMs when using approximate linear solvers: our approach guarantees that, when certain conditions are satisfied, the number of IPM iterations does not increase and that the final solution remains feasible. We also provide practical instantiations of approximate linear solvers that satisfy these conditions for special classes of constraint matrices using randomized linear algebra.

Dexter, Gregory↗

Predicting Volume of Distribution in Humans: Performance of In Silico Methods for a Large Set of Structurally Diverse Clinical Compounds

Volume of distribution at steady state (V D,ss ) is one of the key pharmacokinetic parameters estimated during the drug discovery process. Despite considerable efforts to predict V D,ss , accuracy and choice of prediction methods remain a challenge, with evaluations constrained to a small set (<150) of compounds. To address these issues, a series of in silico methods for predicting human V D,ss directly from structure were evaluated using a large set of clinical compounds. Machine learning (ML) models were built to predict V D,ss directly and to predict input parameters required for mechanistic and empirical V D,ss predictions. In addition, log D, fraction unbound in plasma (fup), and blood-to-plasma partition ratio (BPR) were measured on 254 compounds to estimate the impact of measured data on predictive performance of mechanistic models. Furthermore, the impact of novel methodologies such as measuring partition (Kp) in adipocytes and myocytes (n = 189) on V D,ss predictions was also investigated. In predicting V D,ss directly from chemical structures, both mechanistic and empirical scaling using a combination of predicted rat and dog V D,ss demonstrated comparable performance (62%–71% within 3-fold). The direct ML model outperformed other in silico methods (75% within 3-fold, r 2 = 0.5, AAFE = 2.2) when built from a larger data set. Scaling to human from predicted V D,ss of either rat or dog yielded poor results (<47% within 3-fold). Measured fup and BPR improved performance of mechanistic V D,ss predictions significantly (81% within 3-fold, r 2 = 0.6, AAFE = 2.0). Adipocyte intracellular Kp showed good correlation to the V D,ss but was limited in estimating the compounds with low V D,ss .

59 BASIC BIOLOGICAL SCIENCES↗

Recent progress in atomic-scale controlled plasma processing

Atomic-scale control in plasma processing is becoming increasingly critical for fabricating of advanced semiconductor devices, particularly as the industry shifts toward three-dimensional (3D) architectures and high-aspect-ratio (HAR) structures. This review presents a comprehensive overview of recent developments in atomic-scale controlled plasma processes, organized along two key directions: the hierarchical structure of plasma–surface interactions and the generational evolution of atomic layer processing (ALP) technologies. We examined the gas phase, where molecular design enables selective generation of ions and radicals; the boundary layer, where transport phenomena govern species delivery into nanoscale features, and the surface, where temperature-dependent reactions and cyclic processing determine etching selectivity and precision. Building on this foundation, we outline five generations of ALP—from thermal atomic layer deposition to transport-aware, temporally and structurally decoupled processes—highlighting the increasing sophistication of process control. The review further explores the transition from empirical recipe development to science-based, data-driven methodologies. By integrating quantum-chemical modeling, advanced diagnostics, and machine learning, we demonstrated how predictive models can link plasma species composition to process outcomes, enabling autonomous and adaptive control strategies. Finally, this review discusses the broader societal implications of plasma process innovation through the E4 quartet: energy and resource efficiency, environmental sustainability, evolutionary advancement, and educational promotion. These principles guide the development of sustainable and intelligent atomic-scale manufacturing technologies that are not only technically advanced but also socially responsible.

Ishikawa, Kenji [Nagoya Univ. (Japan)] (ORCID:0000↗

An Entropy-Maximization Approach to Automated Training Set Generation for Interatomic Potentials

Machine learning-based interatomic potentials are currently garnering a lot of attention as they strive to achieve the accuracy of electronic structure methods at the computational cost of empirical potentials. Given their generic functional forms, the transferability of these potentials is highly dependent on the quality of the training set, the generation of which can be highly labor-intensive. Good training sets should at once contain a very diverse set of configurations while avoiding redundancies that incur cost without providing benefits. We formalize these requirements in a local entropy-maximization framework and propose an automated sampling scheme to sample from this objective function. We show that this approach generates much more diverse training sets than unbiased sampling and is competitive with hand-crafted training sets.

74 ATOMIC AND MOLECULAR PHYSICS↗

Training calibration-based counterfactual explainers for deep learning models in medical image analysis

The rapid adoption of artificial intelligence methods in healthcare is coupled with the critical need for techniques to rigorously introspect models and thereby ensure that they behave reliably. This has led to the design of explainable AI techniques that uncover the relationships between discernible data signatures and model predictions. In this context, counterfactual explanations that synthesize small, interpretable changes to a given query while producing desired changes in model predictions have become popular. This under-constrained, inverse problem is vulnerable to introducing irrelevant feature manipulations, particularly when the model’s predictions are not well-calibrated. Hence, in this paper, we propose the TraCE (training calibration-based explainers) technique, which utilizes a novel uncertainty-based interval calibration strategy for reliably synthesizing counterfactuals. Given the wide-spread adoption of machine-learned solutions in radiology, our study focuses on deep models used for identifying anomalies in chest X-ray images. Using rigorous empirical studies, we demonstrate the superiority of TraCE explanations over several state-of-the-art baseline approaches, in terms of several widely adopted evaluation metrics. Our findings show that TraCE can be used to obtain a holistic understanding of deep models by enabling progressive exploration of decision boundaries, to detect shortcuts, and to infer relationships between patient attributes and disease severity.

62 RADIOLOGY AND NUCLEAR MEDICINE↗

Gaussian approximation potential for amorphous Si : H

Hydrogenation of amorphous silicon (a–Si : H) is critical for reducing defect densities, passivating midgap states and surfaces, and improving photoconductivity in silicon-based electro-optical devices. Modeling the atomic-scale structure of this material is critical to understanding these processes, which in turn is needed to describe c–Si/a–Si : H heterojunctions that are at the heart of modern solar cells with world-record efficiency. Density functional theory (DFT) studies achieve the required high accuracy but are limited to moderate system sizes of 100 atoms or so by their high computational cost. Simulations of amorphous materials have been hindered by this high cost because large structural models are required to capture the medium-range order that is characteristic of such materials. Empirical potential models are much faster, but their accuracy is not sufficient to correctly describe the frustrated local structure. Data-driven, machine-learned interatomic potentials have broken this impasse and have been highly successful in describing a variety of amorphous materials in their elemental phase. Here, we extend the Gaussian approximation potential (GAP) for silicon by incorporating the interaction with hydrogen, thereby significantly improving the degree of realism with which amorphous silicon can be modeled. We show that our Si : H GAP enables the simulation of hydrogenated silicon with an accuracy very close to DFT but with computational expense and run times reduced by several orders of magnitude for large structures. Here, we demonstrate the capabilities of the Si : H GAP by creating models of hydrogenated liquid and amorphous silicon and showing that their energies, forces, and stresses are in excellent agreement with DFT results, and their structure as captured by bond and angle distributions are in agreement with both DFT and experiments.

36 MATERIALS SCIENCE↗

Exploring model complexity in machine learned potentials for simulated properties

Abstract Machine learning (ML) enables the development of interatomic potentials with the accuracy of first principles methods while retaining the speed and parallel efficiency of empirical potentials. While ML potentials traditionally use atom-centered descriptors as inputs, different models such as linear regression and neural networks map descriptors to atomic energies and forces. This begs the question: what is the improvement in accuracy due to model complexity irrespective of descriptors? We curate three datasets to investigate this question in terms of ab initio energy and force errors: (1) solid and liquid silicon, (2) gallium nitride, and (3) the superionic conductor Li $$_{10}$$ 10 Ge(PS $$_{6}$$ 6 ) $$_{2}$$ 2 (LGPS). We further investigate how these errors affect simulated properties and verify if the improvement in fitting errors corresponds to measurable improvement in property prediction. By assessing different models, we observe correlations between fitting quantity (e.g. atomic force) error and simulated property error with respect to ab initio values. Graphical abstract

Rohskopf, A. (ORCID:0000000227128296)↗

Predicting Solar Energetic Particles Using SDO/HMI Vector Magnetic Data Products and a Bidirectional LSTM Network

Solar energetic particles (SEPs) are an essential source of space radiation, and are hazardous for humans in space, spacecraft, and technology in general. In this paper, we propose a deep-learning method, specifically a bidirectional long short-term memory (biLSTM) network, to predict if an active region (AR) would produce an SEP event given that (i) the AR will produce an M- or X-class flare and a coronal mass ejection (CME) associated with the flare, or (ii) the AR will produce an M- or X-class flare regardless of whether or not the flare is associated with a CME. The data samples used in this study are collected from the Geostationary Operational Environmental Satellite's X-ray flare catalogs provided by the National Centers for Environmental Information. We select M- and X-class flares with identified ARs in the catalogs for the period between 2010 and 2021, and find the associations of flares, CMEs, and SEPs in the Space Weather Database of Notifications, Knowledge, Information during the same period. Each data sample contains physical parameters collected from the Helioseismic and Magnetic Imager on board the Solar Dynamics Observatory. Experimental results based on different performance metrics demonstrate that the proposed biLSTM network is better than related machine-learning algorithms for the two SEP prediction tasks studied here. We also discuss extensions of our approach for probabilistic forecasting and calibration with empirical evaluation

79 ASTRONOMY AND ASTROPHYSICS↗

The Anatomy of Software Changes and Bugs in Autonomous Operating System

Cyberphysical systems with autonomous functions are complex pieces of software, consisting of many components, some of which implement autonomous functionality and some may use AI or machine learning algorithms. Software bugs in an autonomous system are of particular concern, as they can have catastrophic consequences. However, detailed studies based on empirical data are rare and therefore these bugs are not well understood. This paper aims to contribute towards filling that gap by investigating the software changes and bugs in Autonomy Operating System (AOS) for Unmanned Aircraft Systems (UAS), which consist of 26 components containing about 103,000 lines of code and having a total of 772 bugfixes. Based on the data extracted from the code repository and semi-structured interviews with the developers of AOS, we explore the differences among autonomous software components, components developed using Model-based Software Engineering, and reuse with respect to change proneness, fault proneness, distribution of bugfixes among AOS components and files of these components, and characteristics of bugs of different AOS components. Our results show that the autonomous components were significantly more change prone (measured in number of commits and code churn) and fault prone (measured in bugfixes per KLoC) than non-autonomous components. The distribution of the locations of bugfixes was skewed, both at component and file level (i.e., a small number of components / files contained the majority of bugs). These evidence-based findings provide important insights to researchers and practitioners alike and can be used to efficiently improve the quality and reliability of autonomous systems.

Katerina Goseva-Popstojanova↗

Robust Carbon Dioxide Plume Imaging Using Joint Tomographic Inversion of Seismic Onset Time and Distributed Pressure and Temperature Measurements (Final Report)

We develop and demonstrate rapid and cost-effective methodologies for spatiotemporal tracking of CO2 plumes during geologic sequestration using joint inversion of seismic data and distributed pressure and temperature measurements. Key elements of our methodology are: (a) a computationally efficient approach to pressure and temperature propagation, (b) analysis of time lapse seismic data using a novel ‘seismic onset time’ approach to detect fluid front propagation, and (c) data assimilation and uncertainty assessment via joint inversion of pressure, temperature and time lapse seismic data, and (d) validating the numerical tomographic inversion using a CO2 injection demonstration projects, specifically data collected from the from the Petra Nova Parish Holdings CCUS project in the West Ranch Field, Texas and the Chester-16 reef CO2 injection site in Northern Michigan which is part of the DOE Midwestern Carbon Sequestration Project. The research team is led by Texas A&M University and includes Battelle as a subcontractor with support from Shell, Anadarko, Chevron and JX Nippon. A carbon dioxide (CO2) water-alternating-gas (WAG) pilot was conducted to gain insights into tertiary oil recovery potential via CO2 flood in the West Ranch Field as part of the Petra Nova project, the world’s largest post-combustion CO2 capture and utilization initiative. With a fluvial formation geology and large contrasts in permeability, this is a challenging and novel application of CO2 enhanced oil recovery (EOR). We build a predictive dynamic model of the subsurface that incorporates the multiphase and compositional data acquired during the pilot operation. The calibrated model is used for the carbon dioxide plume imaging. The study began with an initialization of the pilot sector model extracted from a calibrated full-field model. The pilot model calibration follows a two-step hierarchical workflow. First, we performed a large-scale update of the permeability distribution by integrating available bottomhole pressure and multiphase production data. In the second step, local permeability field is fine-tuned using a streamline-based method to match CO2 breakthrough times at the producers. The predictive capability of the calibrated model was verified through two blind validation tests: (1) the model showed good agreement with saturation logs acquired at two observation wells; and (2) the model reproduced the CO2 recovery as a fraction of the injected CO2. The use of seismic onset times has shown great promise for integrating near-continuous seismic surveys for updating geologic models. In this study, we analyze the impact of seismic survey frequency on the onset time approach aiming to extend the application of onset time to infrequent seismic surveys. In addition, we quantitatively examine the nonlinearity of the onset time method and compare it to the commonly used amplitude inversion method. We carry out a sensitivity analysis of seismic survey frequency based on the complete seismic survey data (over 175 surveys) of steam injection in a heavy oil reservoir (Peace River Unit) in Canada. Our results show that an adequate onset time map can be obtained from the infrequent seismic surveys by interpolation between seismic surveys as long as there is no change in the dominant underlying physics between the successive surveys. The study also shows that nonlinearity of the onset time method can be -smaller than that of the amplitude inversion method by several orders of magnitude. Application to the Brugge benchmark case shows that the onset time method obtains comparable permeability update as the traditional seismic amplitude inversion method with faster computation and improved convergence characteristics. We extend the streamline-based data integration approach to incorporate distributed temperature sensor (DTS) data using the concept of thermal tracer travel time. Then, a hierarchical workflow composed of evolutionary and streamline methods is employed to jointly history match the DTS and pressure data. Finally, CO2 saturation and streamline maps are used to visualize the CO2 plume movement during the sequestration process. The hierarchical workflow is applied to a carbon sequestration project in a carbonate reef reservoir within the Northern Niagaran Pinnacle Reef Trend in Michigan, USA. The monitoring data set consists of distributed temperature sensing (DTS) data acquired at the injection well and a monitoring well, flowing bottom-hole pressure data at the injection well, and time-lapse pressure measurements at several locations along the monitoring well. The history matching results indicate that the CO2 movement is mostly restricted to the intended zones of injection which is consistent with an independent warm-back analysis of the temperature data. In addition to employing simulation models and inverse methods for CO2 plume imaging, we also initialized a data-driven technology for detecting inter-well connectivity based on production and pressure data. Our machine-learning framework is built on the statistical recurrent unit (SRU) model and interprets well-based injection/production data into inter-well connectivity without relying on a geologic model. We test it on synthetic and field-scale CO2 EOR projects utilizing the water-alternating-gas (WAG) process. The validation of the proposed data-driven inter-well connectivity assessment is performed using synthetic data from simulation models where inter-well connectivity can be easily measured using the streamline-based flux allocation. The SRU model is shown to offer excellent prediction performance on the synthetic case. Despite significant measurement noise and frequent well shut-ins imposed in the field-scale case, the SRU model offers good prediction accuracy, the overall relative error of the phase production rates at most producers ranges from 10% to 30%. It is shown that the dominant connections identified by the data-driven method and streamline method are in close agreement. Texas A&M University, the lead organization in the project, was primarily responsible for the development of tomographic approaches for CO2 plume mapping in conjunction with distributed pressure, temperature and seismic onset time data. Battelle, as a subcontractor, was primarily responsible for the development of analytical and empirical methods for analyzing transient injection rate and pressure data from point/line sources such as injection and monitoring wells. An additional area of emphasis for Battelle was the use of machine learning for such tasks as inferring reservoir connectivity information from injection-production data, and identifying variable importance for machine learning-based proxy models developed from full-physics simulations. The two organizations also collaborated on the application of the tomographic inversion methodology for a field data set.

02 PETROLEUM↗

Machine learning interatomic potential for silicon-nitride (Si 3 N 4 ) by active learning

Silicon nitride (Si 3 N 4 ) is an extensively used material in the automotive, aerospace, and semiconductor industries. However, its widespread use is in contrast to the scarce availability of reliable interatomic potentials that can be employed to study various aspects of this material on an atomistic scale, particularly its amorphous phase. In this work, we developed a machine learning interatomic potential, using an efficient active learning technique, combined with the Gaussian approximation potential (GAP) method. Our strategy is based on using an inexpensive empirical potential to generate an initial dataset of atomic configurations, for which energies and forces were recalculated with density functional theory (DFT); thereafter, a GAP was trained on these data and an iterative re-training algorithm was used to improve it by learning on-the-fly. When compared to DFT, our potential yielded a mean absolute error of 8 meV/atom in energy calculations for a variety of liquid and amorphous structures and a speed-up of molecular dynamics simulations by 3–4 orders of magnitude, while achieving a first-rate agreement with experimental results. Our potential is publicly available in an open-access repository.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

On the connection between least squares, regularization, and classical shadows

Classical shadows (CS) offer a resource-efficient means to estimate quantum observables, circumventing the need for exhaustive state tomography. Here, we clarify and explore the connection between CS techniques and least squares (LS) and regularized least squares (RLS) methods commonly used in machine learning and data analysis. By formal identification of LS and RLS ``shadows'' completely analogous to those in CS---namely, point estimators calculated from the empirical frequencies of single measurements---we show that both RLS and CS can be viewed as regularizers for the underdetermined regime, replacing the pseudoinverse with invertible alternatives. Through numerical simulations, we evaluate RLS and CS from three distinct angles: the tradeoff in bias and variance, mismatch between the expected and actual measurement distributions, and the interplay between the number of measurements and number of shots per measurement. Compared to CS, RLS attains lower variance at the expense of bias, is robust to distribution mismatch, and is more sensitive to the number of shots for a fixed number of state copies---differences that can be understood from the distinct approaches taken to regularization. Conceptually, our integration of LS, RLS, and CS under a unifying ``shadow'' umbrella aids in advancing the overall picture of CS techniques, while practically our results highlight the tradeoffs intrinsic to these measurement approaches, illuminating the circumstances under which either RLS or CS would be preferred, such as unverified randomness for the former or unbiased estimation for the latter.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Active learning using hybrid surrogate tool life modeling for machining process optimization

Here, this paper describes an active learning approach for part-to-part iterative machining process optimization using a hybrid surrogate tool life model. A probabilistic interpolating tool life model is developed by combining the empirical Taylor-type tool life equation and the model fit error. The probabilistic tool life model is then used to calculate the machining cost per part distribution. The optimal machining parameters are selected using an expected improvement in machining cost per part criterion. The method is validated numerically using experimental results; the results show a median convergence error of 2.2% after three tests over 400 simulations. The method is validated experimentally on two industrial applications for Ti-6Al-4V roughing resulting in a cost per part reduction greater than 23% after two tests. The described method is a robust solution for rapid convergence to optimal machining parameters in an industrial production environment.

Active learning↗