Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data driven optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

An Efficient Numerical Algorithm for Solving Data Driven Feedback Control Problems

The goal of this paper is to solve a class of stochastic optimal control problems numerically, in which the state process is governed by an Itô type stochastic differential equation with control process entering both in the drift and the diffusion, and is observed partially. The optimal control of feedback form is determined based on the available observational data. In this work, we call this type of control problems the data driven feedback control. The computational framework that we introduce to solve such type of problems aims to find the best estimate for the optimal control as a conditional expectation given the observational information. To make our method feasible in providing timely feedback to the controlled system from data, we develop an efficient stochastic optimization algorithm to implement our computational framework.

97 MATHEMATICS AND COMPUTING↗

FluxSat: Long-term Earth Science Data Record (ESDR) for Terrestrial Gross Primary Production (GPP) based on satellite data calibrated with eddy covariance data

Gross primary production (GPP), the amount of carbon dioxide (CO 2 ) assimilated by plants through photosynthesis, is one of the most variable and uncertain components of the global carbon cycle. Global GPP has been estimated with a number of process-based models, data-driven, and hybrid approaches. Dynamic global vegetation models (DGVMs), driven by observed environmental changes, are used for global carbon budget assessments and long-term (climate) prediction. Benchmarking these and other models globally with data-driven GPP estimates is critical for understanding the land sink and ensuring accurate forecasts of the carbon cycle. In addition, global data-driven GPP estimates are crucial for studies of interannual variability, including trends that are linked to mechanisms with large uncertainties, such as the indirect CO 2 fertilization effect related to greening. In response to a community need for a GPP data set that well captures spatio-temporal variability, we developed FluxSat, a data-driven approach that optimizes the use of satellite reflectance data from the NASA MODerate-resolution Imaging Spectroradiometer (MODIS) on the Terra and Aqua satellites, calibrated using ground-based eddy covariance (EC) data. We are enhancing (spatially, higher resolution) and extending FluxSat (in time, with additional sensors) to create a high quality long term GPP Earth System Data Record (ESDR) for use in model benchmarking, carbon cycle modeling, and studies of trends and interannual variability. Our team’s objectives are to: 1. Update and document the current MODIS FluxSat GPP (daily, 0.05o and 0.5o resolutions) products with latest available MODIS and EC data sets; 2. Extend FluxSat GPP record forward in time with the Visible Infrared Imaging Radiometer Suite (VIIRS) on operational weather satellites going forward; 3. Extend FluxSat GPP record backward in time using the Advanced Very High Resolution Radiometer (AVHRR) on weather satellites dating back to 1981; 4. Provide higher spatial resolution MODIS and VIIRS GPP (0.0083o). 5. Thoroughly evaluate all FluxSat products with independent data; and 6. Create a homogenized long-term GPP record spanning 40+ years. We will discuss plans for this long-term data set that is supported through the NASA Making Earth System Data Records for Use in Research Environments (MEaSUREs) program.

gross Primary Production↗

Data-Driven Discovery of Linear Molecular Probes with Optimal Selective Affinity for PFAS in Water

Approaches to tackle the wide and growing variety of highly persistent per- and polyfluoroalkyl substances (PFAS) are of pressing global need because of their detrimental human health effects, such as cancer, birth defects, and hormone imbalance. Sensitive, selective, and easy-to-use real-time sensors to monitor and detect PFAS and sorbents to extract them are critical to meeting government-mandated environmental concentrations. In this work, we combine all-atom molecular dynamics simulations, enhanced sampling, deep representational learning, and Bayesian optimization to perform high-throughput virtual screening for highly sensitive and selective molecular probes. Our molecular design space consists of 3850 linear hydrocarbon chains with varying degrees of halogenation with and without amine- and phosphine-based headgroups. By employing a data-driven search process, we efficiently explore the molecular design space to optimize the sensitivity to perfluorooctanesulfonic acid (PFOS) as a prototypical PFAS analyte and selectivity relative to a sodium dodecyl sulfate (SDS) interferent. We calculate 504 Gibbs free energies of probe-analyte and probe-interferent interactions and identify probes with PFOS association free energies of up to (-ΔG PFOS ) = 9.8 ± 0.2 kJ/mol and selectivities relative to SDS of (-ΔΔG PFOS–SDS ) = 3.1 ± 1.5 kJ/mol. A C 11 Br 23 P(CH 3 ) 2 probe containing 11 backbone brominated carbons and a tertiary phosphine headgroup possesses the most sensitive binding constant to PFOS within the defined search space of K b PFOS = 177.4 ± 12.7, and a semibrominated probe C 5 H 11 C 7 Br 14 N(CH 3 ) 2 containing 12 backbone carbons and a tertiary amine headgroup possesses the highest selectivity relative to SDS of K b PFOS /K b SDS = 4.6 ± 1.7. A retrospective analysis of our data to extract interpretable design rules reveals that the sensitivity of linear hydrogenated probes increases by approximately 1 kJ/mol per C–C bond. The addition or removal of halogen atoms and amine or phosphine headgroups produces nonmonotonic changes in both sensitivity and selectivity with changes to the sensitivity of up to 2.5 kJ/mol. Finally, this work places empirical limitations on the performance of a wide range of linear probes for PFOS detection and offers a generic strategy for high-throughput computational screening to promote selective and sensitive binding.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Data-Driven Generic Turbines for Distributed Wind Modeling, Optimization, and Economic Studies

As distributed energy resources (DER) become less expensive and more popular, utilities, project developers, and customers have an increasing need to model the performance of existing and proposed DER systems. Distributed wind has been shown to have widespread economic potential but is often represented by a simplified model in - or excluded from - DER modeling tools and studies. There is often no economic imperative to extend models and studies to give full consideration to distributed wind. We present a set of data-driven generic turbines derived from 16 years of annual distributed wind market survey data. The proposed methodology can be used to derive generic turbines from separate or updated data sets. Finally, a mixed-integer linear programming approach to optimal distributed wind project sizing is used to demonstrate the generic turbine models. Combined, these models and methods can reduce barriers to considering distributed wind in modeling tools and studies.

Reiman, Andrew P.↗

Optimizing accuracy and efficacy in data-driven materials discovery for the solar production of hydrogen

The production of hydrogen fuels, via water splitting, is of practical relevance for meeting global energy needs and mitigating the environmental consequences of fossil-fuel-based transportation. Water photoelectrolysis has been proposed as a viable approach for generating hydrogen, provided that stable and inexpensive photocatalysts with conversion efficiencies over 10% can be discovered, synthesized at scale, and successfully deployed. While a number of first-principles studies have focused on the data-driven discovery of photocatalysts, in the absence of systematic experimental validation, the success rate of these predictions may be limited. We address this problem by developing a screening procedure with co-validation between experiment and theory to expedite the synthesis, characterization, and testing of the computationally predicted, most desirable materials. Here, starting with 70 150 compounds in the Materials Project database, the proposed protocol yielded 71 candidate photocatalysts, 11 of which were synthesized as single-phase materials. Experiments confirmed hydrogen generation and favorable band alignment for 6 of the 11 compounds, with the most promising ones belonging to the families of alkali and alkaline-earth indates and orthoplumbates. This study shows the accuracy of a nonempirical, Hubbard-corrected density-functional theory method to predict band gaps and band offsets at a fraction of the computational cost of hybrid functionals, and outlines an effective strategy to identify photocatalysts for solar hydrogen generation.

08 HYDROGEN↗

Domain Aware Deep-learning Algorithms Integrated with Scientific-computing Technologies (DADAIST)

This technical report summarized the contribution of the DADAIST project funded by the Data Model Convergence Initiative via the Laboratory Directed Research and Development (LDRD) investments at Pacific Northwest National Laboratory (PNNL). Specifically, we report the development of the NeuroMANCER (Neural Modules with Adaptive Nonlinear Constraints and Efficient Regularizations), a new open-source Scientific Machine Learning library for formulating and solving parametric constrained optimization problems, physics-informed system identification, and parametric optimal control problems. NeuroMANCER is using differentiable programming to combine modern data-driven models and optimization modeling language into a coherent algorithmic and software framework. NeuroMANCER is a Pytorch-based framework and adopts much of its philosophy focused on research and development, rapid prototyping, and streamlined deployment. Strong emphasis is given to extensibility, interoperability with the PyTorch ecosystem, and quick adaptability to custom domain problems. Neuromancer repository contains a comprehensive library of differentiable modules, including custom activation functions, matrix factorizations, deep learning architectures, neural differential equations, differential equation solvers, implicit layers such as iterative solvers, high-level API for symbolic expressions, API for modeling and control of dynamical systems, and extensive set of tutorial code examples in the form of python scripts and jupyter notebooks.

97 MATHEMATICS AND COMPUTING↗

Optimizing aircraft flows at airports using data driven predicted capabilities

A method for safe and efficient use of airport runway capacity includes receiving, at an air traffic control system at an airport, airport data related to movement areas of the airport, time data related to a time period, aircraft data related to a plurality of aircraft expected to operate into and out of the airport during the time period, and environmental data related to environmental conditions predicted for the airport during the time period. The method further includes computing a probability distribution for inter-aircraft spacing by applying the airport data, the time data, the aircraft data, and the environmental data to a trained Bayesian network, producing the probability distribution for the inter-aircraft spacing as an output observation of the trained Bayesian network, and, using the probability distribution and a confidence value, identifying an inter-aircraft spacing value for the plurality of aircraft expected to operate into and out of the airport during the time period.

Sweet, Douglas↗

Multi-level optimization with the koopman operator for data-driven, domain-aware, and dynamic system security

Cyber-Physical Systems (CPSs) like the power grid are critically important but also increasingly vulnerable; ensuring reliable system operation in the face of disruptions is becoming more and more challenging. Multi-Level Optimization (MLO) is a powerful way to model adversarial interactions, which naturally makes it applicable to studying CPS security. However, MLO typically does not address underlying system dynamics, and incorporating nonlinear dynamics is generally infeasible. In this paper, we show how to combine MLO with the Koopman Operator (KO) to remedy this. The KO maps nonlinear dynamics to a lifted space in which those dynamics are linear, thus making it ideal for use with MLO. Moreover, the structure of the KO also provides convenient ways to incorporate domain knowledge into the data-driven process of learning the KO representation of a given system. Here we then demonstrate the use of MLO-KO on a small example problem taken from the power grid domain, discuss the scalability and computational cost of MLO-KO, and identify future research directions for this work.

42 ENGINEERING↗

Sizing ramping reserve using probabilistic solar forecasts: A data-driven method

Ramping products have been introduced or proposed in several U.S. power markets to mitigate the impact of load and renewable uncertainties on market efficiency and reliability. Current methods often rely on historical data to estimate the requirements of ramping products and fail to take into account the effects of the latest weather conditions and their uncertainties, which could lead to overly conservative or insufficient requirements. This study proposes a k-nearest-neighbor-based method to give weather-informed estimates of ramping needs based on short-term probabilistic solar irradiance forecasts. Forecasts from multiple sites are employed in conjunction with principal component analysis to derive numerical classifiers to characterize system-level weather conditions. In addition, we develop a data-driven method to optimize the model parameters in a rolling-forward manner. By using real-world data from the California Independent System Operator, we design two metrics to evaluate method performance: 1) frequency of shortage and 2) oversupply of ramping product. Our proposed method presents advantages in comparison with the baseline and a set of benchmark methods: without compromising system reliability, it reduces system ramping requirements by up to 25%, therefore improving both system reliability and economics.

14 SOLAR ENERGY↗

Dilute Combustion Control Using Spiking Neural Networks

Dilute combustion with exhaust gas recirculation (EGR) in spark-ignition engines presents a cost-effective method for achieving higher levels of engine efficiency. At high levels of EGR, however, cycle-to-cycle variability (CCV) of the combustion process is exacerbated by sporadic occurrences of misfires and partial burns. Previous studies have shown that temporal deterministic patterns emerge at such conditions and certain combustion cycles have a significant influence over future events. Due to the complexity of the combustion process and the nature of CCV, harnessing all the deterministic information for control purposes has remained challenging even with physics based 0-D, 1-D, and high-fidelity computational fluid dynamics (CFD) models. In this study, we present a data-driven approach to optimize the combustion process by controlling CCV adjusting the cycle-to-cycle fuel injection quantity. Readily available data from in-cylinder pressure was used to train a spiking neural network (SNN) which learns the optimal way to manage fuel injection in order to reduce CCV while maintaining acceptable levels of fuel consumption. SNNs are particularly well suited for powertrain control applications due to their ability to be deployed on FPGA-based neuromorphic hardware which are small, inexpensive, and have a low power demand. The high-performance computing (HPC) resources of Oak Ridge National Laboratory were used to run an evolutionary-based training approach for choosing the best SNN configuration that minimizes the size of the network while achieving the desired goal. The neuromorphic hardware with the optimized SNN deployed was connected to the rapid prototyping engine control system for real-time control implementation and tested on a single cylinder version of a GM LNF 4-cylinder engine. The results show a significant reduction of CCV with a small percentage of additional fuel used to stabilize the charge.

33 ADVANCED PROPULSION SYSTEMS↗

Data-driven analysis to understand GPU hardware resource usage of optimizations

With heterogeneous systems, the number of GPUs per chip increases to provide computational capabilities for solving science at a nanoscopic scale. However, low utilization for single GPUs defies the need to invest more money in expensive accelerators. Although related work develops optimizations to improve application performance, none studies how these optimizations impact hardware resource usage or average GPU utilization. Here, this paper takes a data-driven analysis approach in addressing this gap by (1) characterizing how hardware resource usage affects device utilization, execution time, or both, (2) presenting a multiobjective metric to identify important application-device interactions that can be optimized to improve device utilization and application performance jointly, (3) studying hardware resource usage behaviors of several optimizations for a benchmark application, and finally (4) identifying optimization opportunities for several scientific proxy applications based on their hardware resource usage behaviors. Furthermore, we demonstrate the applicability of our methodology by applying the identified optimizations to a proxy application, which improves the execution time, device utilization, and power consumption by up to 29.6%, 5.3% and 26.5% respectively.

Computer science↗

Integrated data-driven and experimental approaches to accelerate lead optimization targeting SARS-CoV- 2 main protease

Identification of potential therapeutic candidates can be expedited by integrating computational modeling with domain aware machine learning (ML) approaches followed by experimental validation. Generative deep learning models have been recently developed that can generate thousands of new candidates, but their physiochemical properties are typically not optimized. Using our deep learning models and a scaffold as a starting point, we generated tens of thousands of compounds for SARS-CoV-2 M pro that preserve the core scaffold. Here we utilized and implemented several computational tools such as structural alert and toxicity analysis, high throughput virtual screening, ML-based 3D quantitative structure–activity relationships, multi-parameter optimization, and graph neural networks on libraries of generated candidates to predict biological activity and binding affinity a priori. From these collective computational results, eight promising candidates were identified and tested experimentally using Native Mass Spectrometry (MS) and FRET-based functional assays. Two compounds, with quinazoline-2-thiol and acetylpiperidine core moiety showed IC 50 values in the low micromolar range: 2.95±0.0017 µM and 3.41±0.0015 µM, respectively. The molecular dynamics simulations further highlight that binding of these compounds results in allosteric modulations in the chain B and the interface domains of the M pro . The key fragments from these top hits can be used as input for closed loop lead optimization in the integrated pipeline.

60 APPLIED LIFE SCIENCES↗

Learning-Accelerated ADMM for Distributed DC Optimal Power Flow

We suggest a novel data-driven method to accelerate the convergence of Alternating Direction Method of Multipliers (ADMM) for solving distributed DC optimal power flow (DC-OPF) where lines are shared between independent network partitions. Using previous observations of ADMM trajectories for a given system under varying load, the method trains a recurrent neural network (RNN) to predict the converged values of dual and consensus variables. Given a new realization of system load, a small number of initial ADMM iterations is taken as input to infer the converged values and directly inject them into the iteration. We empirically demonstrate that the online injection of these values into the ADMM iteration accelerates convergence by a significant factor for partitioned 14-, 118-and 2848-bus test systems under differing load scenarios. The proposed method has several advantages: it maintains the security of private decision variables inherent in consensus ADMM; inference is fast and so may be used in online settings; RNN-generated predictions can dramatically improve time to convergence but, by construction, can never result in infeasible ADMM subproblems; it can be easily integrated into existing software implementations. While we focus on the ADMM formulation of distributed DC-OPF in this paper, the ideas presented are naturally extended to other distributed optimization problems.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Learning-Accelerated ADMM for Distributed DC Optimal Power Flow

We propose a novel data-driven method to accelerate the convergence of Alternating Direction Method of Multipliers (ADMM) for solving distributed DC optimal power flow (DC-OPF) where lines are shared between independent network partitions. Using previous observations of ADMM trajectories for a given system under varying load, the method trains a recurrent neural network (RNN) to predict the converged values of dual and consensus variables. Given a new realization of system load, a small number of initial ADMM iterations is taken as input to infer the converged values and directly inject them into the iteration. We empirically demonstrate that the online injection of these values into the ADMM iteration accelerates convergence by a significant factor for partitioned 14-, 118- and 2848-bus test systems under differing load scenarios. The proposed method has several advantages: it maintains the security of private decision variables inherent in consensus ADMM; inference is fast and so may be used in online settings; RNN-generated predictions can dramatically improve time to convergence but, by construction, can never result in infeasible ADMM subproblems; it can be easily integrated into existing software implementations. While we focus on the ADMM formulation of distributed DC-OPF in this paper, the ideas presented are naturally extended to other distributed optimization problems.

alternating direction method of multipliers↗