Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Data driven”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Data-Conforming Data-Driven Control: Avoiding Premature Generalizations Beyond Data

Data-driven and adaptive control approaches face the problem of introducing sudden distributional shifts beyond the distribution of data encountered during learning. Therefore, they are prone to invalidating the very assumptions used in their own construction. This is due to the linearity of the underlying system, inherently assumed and formulated in most data-driven control approaches, which may falsely generalize the behavior of the system beyond the behavior experienced in the data. This article seeks to mitigate these problems by enforcing consistency of the newly designed closed-loop systems with data and slowing down any distributional shifts in the joint state-input space. This is achieved through incorporating affine regularization terms and linear matrix inequality constraints to data-driven approaches, resulting in convex semi-definite programs that can be efficiently solved by standard software packages. We discuss the optimality conditions of these programs and then conclude this article with a numerical example that further highlights the problem of premature generalization beyond data and shows the effectiveness of our proposed approaches in enhancing the safety of data-driven control methods.

97 MATHEMATICS AND COMPUTING

Creating Accurate Methane Emission Inventories through Data-Driven Airborne Survey Strategies

Because natural gas emits less carbon than other fossil fuels, it holds promise as a green energy transition fuel. However, the overall carbon footprint of natural gas is significantly elevated by methane emissions that occur during its production and transmission (Cusworth et al. 2022). Methane “super-emitters,” while comprising only about 1% of sites, are responsible for the majority of oil- and gas-sourced methane emissions, making their detection and mitigation critical in reducing the climate impact of natural gas and in meeting national and global sustainability goals (Sherwin et al. 2024). Yet, despite advancements in detection, significant uncertainties remain regarding the size, frequency, and duration distributions of methane emissions (e.g., Frankenberg et al. 2016, Cusworth et al. 2022, Chen, Sherwin et al. 2022, Conrad et al. 2023, Johnson et al. 2023, Sherwin et al. 2024) underscoring the need for comprehensive emissions inventories segmented by basin across the US. Airborne surveys are well-suited for collecting data to build these comprehensive, basin-level inventories because they allow for extensive spatial coverage, and have the spatial resolution, and the sensitivity to pinpoint individual methane sources. As remote sensing technologies enable rapid basin-scale surveys, it is imperative to establish scientifically and statistically robust standards to generate reliable and actionable emissions inventories. Recent work has shown that differences in airborne sampling strategies, detection technologies, and analysis can lead to large differences between survey conclusions if not correctly accounted for (Chen et al. 2024). This elevates the importance of incorporating proper sampling and analysis techniques when designing a methane emissions monitoring campaign to produce accurate results and facilitate cross-study comparisons. In this paper, we describe a survey strategy designed using the latest conclusions from the literature to align results from different aerial surveys. We identify several sampling and analysis principles, including large sample sizes, balanced sampling across oil and gas production, careful survey area definition, and a unified protocol for analysis, to be vital to producing an unbiased estimate of basin-scale emissions. We present results from a Department of Energy-funded project that deployed this survey strategy in two understudied oil and gas- producing regions in the United States: the Haynesville Basin in Texas and Louisiana, and the Woodford Shale in the Anadarko Basin in Oklahoma.

03 NATURAL GAS

Photospheric Current Spikes and Their Possible Association with Flares - Results from an HMI Data Driven Model

A data driven, near photospheric magnetohydrodynamic model predicts spikes in the horizontal current density, and associated resistive heating rate per unit volume Q. The spikes appear as increases by orders of magnitude above background values in neutral line regions (NLRs) of active regions (ARs). The largest spikes typically occur a few hours to a few days prior to M or X flares. The spikes correspond to large vertical derivatives of the horizontal magnetic field. The model takes as input the photospheric magnetic field observed by the Helioseismic & Magnetic Imager (HMI) on the Solar Dynamics Observatory (SDO) satellite. This 2.5 D field is used to determine an analytic expression for a 3 D magnetic field, from which the current density, vector potential, and electric field are computed in every AR pixel for 14 ARs. The field is not assumed to be force-free. The spurious 6, 12, and 24 hour Doppler periods due to SDO orbital motion are filtered out of the time series of the HMI magnetic field for each pixel using a band pass filter. The subset of spikes analyzed at the pixel level are found to occur on HMI and granulation scales of 1 arcsec and 12 minutes. Spikes are found in ARs with and without M or X flares, and outside as well as inside NLRs, but the largest spikes are localized in the NLRs of ARs with M or X flares. The energy to drive the heating associated with the largest current spikes comes from bulk flow kinetic energy, not the electromagnetic field, and the current density is highly non-force free. The results suggest that, in combination with the model, HMI is revealing strong, convection driven, non-force free heating events on granulation scales, and that it is plausible these events are correlated with subsequent M or X flares. More and longer time series need to be analyzed to determine if such a correlation exists. Above an AR dependent threshold value of Q, the number of events N(Q) with heating rates greater than or equal to Q obeys a scale invariant power law distribution for each AR given by N(Q) varies Q(sup -s), where 0.40 less than or equal to S less than or equal to 0.53, with a mean and standard deviation across the 14 ARs of 0.47 and 0.045, showing there is little variation of S from one AR to another. These properties of N(Q) are in close agreement with those of the distribution N(E) for the total energy E of solar flares, determined from observations to be N(E) = constant x E(sup -alpha). From observations of nanoflares in the 0.7 to 4 MK range, and from observations of flares in hard X-rays, it is found that 0.51 less than or equal to alpha less than or equal to 0.57, and 0.4 less than or equal to alpha less than or equal to 0.6, respectively (Crosby et al. 1993, Sol. Phys., 143, 275; Aschwanden & Parnell 2002, ApJ, 572, 1048). Observations also show that, as is found here for the exponent S, there is little variation of alpha with AR (Wheatland 2000, ApJ, 532, 1209), indicating N(E) and N(Q) are largely independent of individual properties of ARs such as area, total magnetic flux, and distribution of current density (i.e. non-potentiality). Therefore the power law scaling of the photospheric heating rate Q computed here on granulation scales is essentially identical to that found for coronal observations of flare energies on scales 1-2 orders of magnitude larger. This suggests the physical mechanisms that cause Q and coronal flares are closely related. It seems likely that Q is the signature of a magnetic reconnection process in an energy range and volume orders of magnitude smaller than those of flares. In this context, at least the larger spikes in Q might be signatures of UV photospheric or lower chromospheric bombs in which plasma is heated to temperatures approximately 10(exp -5) K (Peter et al. 2014, Science 346, 1255726; Judge 2015, ApJ, 808, 116). In addition, lattice based avalanche simulations of flare energy release predict 0.4 less than or equal to alpha less than or equal to 0.5, while analytic, fractal-diffusive self-organized criticality models predict 0.4 less than or equal to alpha less than or equal to 0.67, in excellent agreement with observations, and the results presented here (Aschwanden & Parnell 2002, ApJ, 572, 1048; Aschwanden 2012, A&A, 539, A2; Aschwanden 2013, in "Self Organized Criticality Systems"; Aschwanden et al. 2016, SSR, 198, 47).

Goodman, Michael

Estimation of Terrestrial Global Gross Primary Production (GPP) with Satellite Data-Driven Models and Eddy Covariance Flux Data

We estimate global terrestrial gross primary production (GPP) based on models that use satellite data within a simplified light-use efficiency framework that does not rely upon other meteorological inputs. Satellite-based geometry-adjusted reflectances are from the MODerate-resolution Imaging Spectroradiometer (MODIS) and provide information about vegetation structure and chlorophyll content at both high temporal (daily to monthly) and spatial (1 km) resolution. We use satellite-derived solar-induced fluorescence (SIF) to identify regions of high productivity crops and also evaluate the use of downscaled SIF to estimate GPP. We calibrate a set of our satellite-based models with GPP estimates from a subset of distributed eddy covariance flux towers (FLUXNET 2015). The results of the trained models are evaluated using an independent subset of FLUXNET 2015 GPP data. We show that variations in light-use efficiency (LUE) with incident PAR are important and can be easily incorporated into the models. Unlike many LUE-based models, our satellite-based GPP estimates do not use an explicit parameterization of LUE that reduces its value from the potential maximum under limiting conditions such as temperature and water stress. Even without the parameterized downward regulation, our simplified models are shown to perform as well as or better than state-of-the-art satellite data-driven products that incorporate such parameterizations. A significant fraction of both spatial and temporal variability in GPP across plant functional types can be accounted for using our satellite-based models. Our results provide an annual GPP value of 140 Pg C year 1 for 2007 that is within the range of a compilation of observation-based, model, and hybrid results, but is higher than some previous satellite observation-based estimates

CO2

New Data-Driven Estimation of Terrestrial CO2 Fluxes in Asia Using a Standardized Database of Eddy Covariance Measurements, Remote Sensing Data, and Support Vector Regression

The lack of a standardized database of eddy covariance observations has been an obstacle for data-driven estimation of terrestrial carbon dioxide fluxes in Asia. In this study, we developed such a standardized database using 54 sites from various databases by applying consistent postprocessing for data-driven estimation of gross primary productivity (GPP) and net ecosystem carbon dioxide exchange (NEE). Data-driven estimation was conducted by using a machine learning algorithm: support vector regression (SVR), with remote sensing data for 2000 to 2015 period. Site-level evaluation of the estimated carbon dioxide fluxes shows that although performance varies in different vegetation and climate classifications, GPP and NEE at 8 days are reproduced (e.g., r (exp 2) =0.73 and 0.42 for 8 day GPP and NEE). Evaluation of spatially estimated GPP with Global Ozone Monitoring Experiment 2 sensor-based Sun-induced chlorophyll fluorescence shows that monthly GPP variations at subcontinental scale were reproduced by SVR (r (exp 2)=1.00, 0.94, 0.91, and 0.89 for Siberia, East Asia, South Asia, and Southeast Asia, respectively). Evaluation of spatially estimated NEE with net atmosphere-land carbon dioxide fluxes of Greenhouse Gases Observing Satellite (GOSAT) Level 4A product shows that monthly variations of these data were consistent in Siberia and East Asia; meanwhile, inconsistency was found in South Asia and Southeast Asia. Furthermore, differences in the land carbon dioxide fluxes from SVR-NEE and GOSAT Level 4A were partially explained by accounting for the differences in the definition of land carbon dioxide fluxes. These data-driven estimates can provide a new opportunity to assess carbon dioxide fluxes in Asia and evaluate and constrain terrestrial ecosystem models.

chlorophyll fluorescence

A data-driven framework for predicting machining stability: employing simulated data, operational modal analysis, and enhanced transfer learning

Chatter, a self-excited vibration phenomenon, presents a significant challenge in machining operations, particularly in high-speed milling, where it can degrade tool life, reduce material removal efficiency, and compromise workpiece quality. Addressing this challenge requires a reliable predictive model that can accommodate the complex dynamics of various machining scenarios. This study introduces a novel, data-driven approach to predicting machining stability, leveraging over 140,000 simulated datasets and employing advanced techniques such as operational modal analysis (OMA), enhanced transfer learning (TL), and receptance coupling substructure analysis (RCSA). By integrating these methodologies, the framework effectively classifies and predicts chatter across diverse operational modes, achieving robust and accurate outcomes. Our model utilizes a Random Forest (RF) classifier trained with the comprehensive dataset, which demonstrates substantial improvements in both predictive accuracy and robustness. Specifically, the RF model achieved an accuracy rate of 85%, an area under the curve (AUC) of 0.90, and an F1 score of 0.88, underscoring its capability to adapt to varying machining configurations. These results highlight the framework’s potential to enhance operational efficiency and machining quality by providing reliable chatter predictions across a broad range of machining parameters. In conclusion, this research thus offers a significant advancement in predictive maintenance for machining processes, enabling more stable and efficient manufacturing operations.

42 ENGINEERING

Predicting Li-Ion Battery Capacity Fade Using Early-Life Data and a Hybrid Data-Driven Gaussian Process-Bayesian Regression Approach

Accurately predicting Li-ion battery capacity trajectories using early-life data can dramatically improve battery-life understandings and be used to rapidly evaluate design/cost/performance trade-offs when developing new battery materials. Accurate early-life predictions enable researchers to quickly iterate over cell designs and material precursor properties without consistently cycling cells to failure. To this end, we present a toolbox that uses a combined Gaussian Process and Bayesian regression approach that capitalizes on signals other than just capacity (e.g., dQ/dV, voltage drops) to rapidly predict capacity-fade trajectories. The prediction tool uses Bayesian regression to fit functional forms, e.g., power law, sigmoids, etc., to predict capacity-fade dynamics. By fitting functional forms, the capacity fade can be interrogated at any point in the future, allowing for early cell-failure prediction. Additionally, Bayesian regression allows for accurate uncertainty estimates that account for cell-to-cell variability (aleatoric uncertainty) and the lack of observation data (epistemic uncertainty). By only using early cycle data to predict the capacity fade trajectory, uncertainty bounds at end-of-life can be extremely large. The large uncertainty bounds are further exacerbated because there is no systematic way to define the prior distribution of the functional forms' parameters. We improve our the predicted trajectory confidence interval of our predicted trajectory using two methods. First, we shows that a small amount of held-out cycling data is sufficientuse some train cells, that have been cycled to failure to derive information regarding the appropriate prior distributions for the functional forms' parameters of the functional form, effectively leading to data-driven priors.. We propose constructing the data-driven priors by first running a Bayesian regression starting with uninformed priors to generate intermediate cell-specific posterior parameter distributions. These posterior distributions are combined using a Ggaussian mixture model for each parameter to create the data-driven priors. These mixture models serve as the data-driven prior distributions for the parameters for. Second, we derive multiple features, e.g., C_dchg 0.5 DoD 0.5, log (|mean(dQ/dV_(w_3-w_0 ) (V)|), etc., from the train cellsheld-out cycling data, identify which the features are that best predicting capacity at early/mid-life cycles, and then create Ggaussian process regression models that are used for predicting capacity at early/mid-life cycles for the test cells (see blue dots with error bars in Fig 1b). Finally, these predicted data-points are used in addition to the actual early cycle data capacity fade to construct the Bayesian regression trajectory for the test cell s. Notably. We note that these two methods are complementary and can be combined with each other. We evaluate the performance of our proposed method on an testing open-source dataset from Iowa State University and Iowa Lakes Community College (ISU-ILCC). This dataset comprises of 251 nickel-manganese-cobalt/graphite Lithium-ion cells that are cycled under 63 different conditions. We compute the mean average percentage error (MAPE) and negative log predictive density (NLPD) to quantify the efficacy of our method. Our initial findings suggest that, when only few observations are available, for test cells, when using only Bayesian regression with uninformed priors, a power law functional provides the most accurate predictions. with very few data points. However, asHowever, a the number of data points increases, a twin sigmoidal function becomes more accurate as the number of observations further increases. We also find that using as little as 10% of the data set towards generating data-driven priors can lead to significant improvement in prediction accuracy when using early cycle data. Lastly, we found that augmenting early-cycle data with Gaussian process-predicted capacity data for Bayesian regression greatly improves the prediction accuracy. We will present a comprehensive comparison of our methods to other methods available in the literature and apply this method to additional battery datasets.

42 ENGINEERING

Influence of initial conditions on data-driven model identification and information entropy for ideal mhd problems

Data-driven methods of model identification are able to discern governing dynamics of a system from data. Such methods are well suited to help us learn about systems with unpredictable evolution or systems with ambiguous governing dynamics given our current understanding. Many plasma problems of interest fall into these categories as there are a wide range of models that exist, however each model is only useful in a certain regime and often limited by computational complexity. To ensure data-driven methods align with theory, they must be consistent and predictable when acting on data whose governing dynamics are known. Weak Sparse Identification of Nonlinear Dynamics (WSINDy) is a recently developed data-driven method that has shown promise in learning governing dynamics from data with high noise levels [1]. This work examines how WSINDy acts on ideal MHD test problems as the initial conditions are varied and specifies limiting requirements for successful equation identification. Furthermore, it is hard to recover the governing dynamics from data that emphasize a single dominant behavior. In these low information cases, Shannon information entropy is able to pick up on the redundancies in the data that affect recoverability.

97 MATHEMATICS AND COMPUTING

The Deep-Time Digital Earth program: data-driven discovery in geosciences

Current barriers hindering data-driven discoveries in deep-time Earth (DE) include: substantial volumes of DE data are not digitized; many DE databases do not adhere to FAIR (findable, accessible, interoperable and reusable) principles; we lack a systematic knowledge graph for DE; existing DE databases are geographically heterogeneous; a significant fraction of DE data is not in open-access formats; tailored tools are needed. These challenges motivate the Deep-Time Digital Earth (DDE) program initiated by the International Union of Geological Sciences and developed in cooperation with national geological surveys, professional associations, academic institutions and scientists around the world. DDE’s mission is to build on previous research to develop a systematic DE knowledge graph, a FAIR data infrastructure that links existing databases and makes dark data visible, and tailored tools for DE data, which are universally accessible. DDE aims to harmonize DE data, share global geoscience knowledge and facilitate data-driven discovery in the understanding of Earth’s evolution.

Chengshan Wang

A cross-dimensional analysis of data-driven short-term load forecasting methods with large-scale smart meter data

Electricity load forecasting is essential to utility operation and power grid stability. A wide spectrum of data-driven methods, ranging from linear regression models to more recent deep learning models have been adopted to forecast electric load over the years. However, there still lacks a holistic evaluation of the applicability of conventional statistical and machine learning based algorithms with respect to different temporal and spatial scopes, computational requirements, and sensitivity of model-tuning. Enabled by a large-scale electricity load profile dataset of over 40,000 residential customers in a utility region, we conducted a cross-dimensional analysis of data-driven load forecasting methods. Three regression-based and seven deep learning algorithms with different model configurations were evaluated in terms of their overall and peak load prediction accuracy, and training burdens, across spatial aggregation levels ranging from the transformer, feeder, substation, to neighborhood. We found, first, the load forecasting accuracy is constrained by a predictability boundary, influenced by the forecasting horizon and spatial aggregation level. Specifically, RandomForest, XGBoost, TFT, TSMixer, and TiDE models achieved less than 10 % prediction error for up to 96-h ahead forecasting for district, substation, and feeder levels, while other models struggle at long-horizon predictions; Second, for winter and summer peak load dates, most models were able to predict the peak demand timing within ± 1 h, but the prediction percentage error varied by models, with TFT and TiDE models being the top performers; Third, models with similar prediction accuracy can differ in training burden by an order of magnitude. Therefore, choosing model configurations that balance prediction performance and computational resource is an important practical consideration for large-scale deployment of the machine learning based load forecasting. The outcome of this study can guide researchers and practitioners to choose the proper load forecasting algorithms based on their problem scope, required accuracy, and available resources. The predictability boundary can serve as a benchmark for electricity load forecasting problems with new algorithms and datasets.

Li, Han

A Comparative Study of Physics‐Informed and Data‐Driven Neural Networks for Compound Flood Simulation at River‐Ocean Interfaces: A Case Study of Hurricane Irene

Simulating compound flooding (CF) at the river-ocean interface within large-scale Earth System Models (ESMs) presents significant challenges due to complex interactions between river discharge, storm surge, and tides. This study assesses the comparative advantages of physics-informed and data-driven machine learning (ML) approaches for enhancing local ESM performance. We systematically compare data-driven neural network models (i.e., CNNs, U-Net, Long Short-Term Memory (LSTM), Gated Recurrent Unit), and physics-informed neural network (PINN) models, including vanilla PINN and a finite-difference-based PINN (FD-PINN). Specifically, FD-PINN is introduced to enhance computational efficiency, accelerating vanilla PINNs by ∼6.5 times while improving accuracy. To enhance data-driven model training, a new data-generation approach is developed to sample historical fluvial and coastal flood events, which ensures a robust data set for extreme event prediction. The models are evaluated using a realistic one-dimensional river domain extracted from an ESM's river mesh and the Hurricane Irene event as an independent test case. Results show that FD-PINN achieves accurate predictions with significantly reduced computational costs relative to vanilla PINNs. Among data-driven models, the best overall performance is achieved by a CNN-LSTM hybrid, which balances accuracy and efficiency. While a fully connected CNN (CNN-FC) provides the best accuracy, it incurs high computational cost. Architectures lacking strong temporal modeling tend to underperform on unseen events. These findings highlight the importance of sequence-aware designs for robust generalization. This study reveals the trade-offs between physics-informed and data-driven models and proposes an adaptive hybrid framework for integrating ML into ESMs to enhance local flood simulations.

Earth Systems Modeling

Data-Driven Modeling and Correction of Vehicle Dynamics

We develop a data-driven framework for learning and correcting nonautonomous vehicle dynamics. Physics-based vehicle models are often simplified for tractability and therefore exhibit inherent model-form uncertainty, motivating the need for data-driven correction. Moreover, nonautonomous dynamics are governed by time-dependent control inputs, which pose challenges in learning predictive models directly from temporal snapshot data. To address these, we reformulate the vehicle dynamics via a local parameterization of the time-dependent inputs, yielding a modified system composed ofa sequence of local parametric dynamical systems. Here, we approximate these parametric systems using two complementary approaches. First, we employ the dimension reduction and interpolation in parameter space (DRIPS) methodology to construct efficient linear surrogate models, equipped with lifted observable spaces and manifold-based operator interpolation. This enables data-efficient learning of vehicle models whose dynamics admit accurate linear representations in the lifted spaces. Second, for more strongly nonlinear systems, we employ flow map learning (FML), a deep neural network (DNN) approach that approximates the parametric evolution map without requiring special treatment of nonlinearities. We further extend FML with a transfer-learning-based model correction procedure, enabling the correction of misspecified prior models using only a sparse set of high-fidelity or experimental measurements, without assuming a prescribed form for the correction term. Through a suite of numerical experiments on unicycle, simplified bicycle, and slip-based bicycle models, we demonstrate that DRIPS offers robust and highly data-efficient learning of nonautonomous vehicle dynamics, while FML provides expressive nonlinear modeling and effective correction of model-form errors under severe data scarcity.

data-driven modeling

Enabling Mission Flexibility to Battery Driven Deep Space Endeavors With Generalized Battery-Health-Monitoring Using Physics-Based and Data-Driven Reduced-Order Models

The needs and requirements for an electrochemical energy storage for deep space exploration is well explored. It is often understood that different mission sites and environmental conditions require different battery chemistries or technologies. Additionally, various engineering solutions are deployed to overcome specific chemical challenges. One often overlooked need is the “health” monitoring of an electrochemical storage system. The term generalized health monitoring, as envisioned in this work, refers to the monitoring of various aspects such as electrode health, electrolyte health, reaction pathway health, cooling system health, sensor health, and BMS health [1]. Generalized health monitoring allows mission leads, engineers, and scientists to incorporate flexibility in mission designs, make on-the-fly mission changes, and extend the duration of science missions. Moreover, it enables automation and data-driven decision-making without compromising safety and performance. Recently, our group developed a hierarchy of thermal reduced-order models (TROM) by combining a physics-based modeling approach and data-driven model reduction techniques applied to flight data [2]. The resulting TROMs were found to be not only accurate but also identifiable from the flight data. Consequently, the coefficient of variance of the model parameters is small over the course of hundreds of flights, allowing for monitoring the parameter evolution trajectories as the battery ages and degrades. These parameters constitute the metrics of the generalized health of a battery. Monitoring their evolution allows such models to be used for anomaly detection and prognostics, improving early detection of abnormal behavior and thus enabling timely maintenance, longer battery life, and enhanced battery safety. For this presentation, the practicality of the thermal model will be validated on a pack of 14cells under various topology configurations such as 1S14P, 2P7S, 7S2P, and 1P14S. It is well known that manufacturing and non-uniform aging lead to variability in the performance of a cell, which is exacerbated by cell balancing during active load. Additionally, in extreme scenarios, the paramount objective is to complete the mission, regardless of the stresses on the battery. Topology-induced balancing issues further stress the battery. The goal of this study is to determine if the noise (identifiability) in the reduced-order thermal model parameters is sensitive to topology, cell spacing, cooling strategy, and manufacturing or age variability. The variability in cells is considered by assuming a multimodal distribution for microscopic parameters of a cell (such as porosity, tortuosity, reaction kinetics, volumetric thermal conductivity, and volumetric heat capacity). The compounded effect of manufacturing variability, topological selection, cooling strategies, and cell balancing ages each cell in a battery differently. The study aims to clarify whether the challenge in extracting maximum information depends on the minimum number of sensors or models used for data extraction.

Automation

Convergence Analysis for an Online Data-Driven Feedback Control Algorithm

This paper presents convergence analysis of a novel data-driven feedback control algorithm designed for generating online controls based on partial noisy observational data. The algorithm comprises a particle filter-enabled state estimation component, estimating the controlled system’s state via indirect observations, alongside an efficient stochastic maximum principle-type optimal control solver. By integrating weak convergence techniques for the particle filter with convergence analysis for the stochastic maximum principle control solver, we derive a weak convergence result for the optimization procedure in search of optimal data-driven feedback control. Numerical experiments are performed to validate the theoretical findings.

97 MATHEMATICS AND COMPUTING

In Search of Data-Driven Improvements to RANS Models Applied to Separated Flows

The goal of this work is to improve the capability of Reynolds-averaged Navier-Stokes turbulence models for separated flows using data-driven enhancements. The resulting model should be “universal” in the sense that it can be used by anyone and applied to as many flows as possible without concern for unusual or detrimental behavior. At worst, the data-driven corrections should not degrade the accuracy of the baseline model (in this case the Spalart-Allmaras one-equation model), while preserving the Galilean invariance and similar theoretical qualities of the original model. In the literature, most current data-driven improvements to turbulence models are only applicable to very similar types of cases as those used to train the model for a specific class of flows. In this work, the impact of using a wide array of cases in the machine-learning training is described. Unwanted behaviors from trained neural networks are examined, and possible mitigation strategies are proposed. However, to date, consistent and broadly applicable data-driven improvements for separated flows have not been achieved.

turbulence modeling

General Purpose Data-Driven Monitoring for Space Operations

As modern space propulsion and exploration systems improve in capability and efficiency, their designs are becoming increasingly sophisticated and complex. Determining the health state of these systems, using traditional parameter limit checking, model-based, or rule-based methods, is becoming more difficult as the number of sensors and component interactions grow. Data-driven monitoring techniques have been developed to address these issues by analyzing system operations data to automatically characterize normal system behavior. System health can be monitored by comparing real-time operating data with these nominal characterizations, providing detection of anomalous data signatures indicative of system faults or failures. The Inductive Monitoring System (IMS) is a data-driven system health monitoring software tool that has been successfully applied to several aerospace applications. IMS uses a data mining technique called clustering to analyze archived system data and characterize normal interactions between parameters. The scope of IMS based data-driven monitoring applications continues to expand with current development activities. Successful IMS deployment in the International Space Station (ISS) flight control room to monitor ISS attitude control systems has led to applications in other ISS flight control disciplines, such as thermal control. It has also generated interest in data-driven monitoring capability for Constellation, NASA's program to replace the Space Shuttle with new launch vehicles and spacecraft capable of returning astronauts to the moon, and then on to Mars. Several projects are currently underway to evaluate and mature the IMS technology and complementary tools for use in the Constellation program. These include an experiment on board the Air Force TacSat-3 satellite, and ground systems monitoring for NASA's Ares I-X and Ares I launch vehicles. The TacSat-3 Vehicle System Management (TVSM) project is a software experiment to integrate fault and anomaly detection algorithms and diagnosis tools with executive and adaptive planning functions contained in the flight software on-board the Air Force Research Laboratory TacSat-3 satellite. The TVSM software package will be uploaded after launch to monitor spacecraft subsystems such as power and guidance, navigation, and control (GN&C). It will analyze data in real-time to demonstrate detection of faults and unusual conditions, diagnose problems, and react to threats to spacecraft health and mission goals. The experiment will demonstrate the feasibility and effectiveness of integrated system health management (ISHM) technologies with both ground and on-board experiments.

Iverson, David L.

Data-driven results for light-quark connected and strange-plus-disconnected hadronic g − 2 short- and long-distance windows

A key issue affecting the attempt to reduce the uncertainty on the Standard Model prediction for the muon anomalous magnetic moment is the current discrepancy between lattice-QCD and data-driven results for the hadronic vacuum polarization. Progress on this issue benefits from precise data-driven determinations of the isospin-limit light-quark-connected (lqc) and strange-plus-light-quark-disconnected ( s + lqd ) components of the related RBC/UKQCD windows. In this paper, using a strategy employed previously for the intermediate window, we provide data-driven results for the lqc and s + lqd components of the short- and long-distance RBC/UKQCD windows. Comparing these results with those from the lattice, we find significant discrepancies in the lqc parts but good agreement for the s + lqd components. We also explore the impact of recent CMD- 3 e + e − → π + π − cross section results, demonstrating that an upward shift in the ρ -peak region of the type seen in the CMD-3 data serves to eliminate the discrepancies for the lqc components without compromising the good agreement between lattice and data-driven s + lqd results. Published by the American Physical Society 2025

Benton, Genessa (ORCID:0009000515763654)