Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “ensemble optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

maestro

MÆSTRO stands for Multi-fidelity Adaptive Ensemble Stochastic Trust Region Optimization and it is a plug n play derivate fee stochastic optimization solver. The problem being considered in MÆSTRO involves fitting Monte Carlo simulations that describe complex phenomena to experiments. This is done by finding parameters of the resource intensive and noisy simulation that yield the least squares objective function value to the noisy experimental data. This problem is solved using a stochastic trust-region optimization algorithm where in each iteration, a local approximation of the simulation signal and of the simulation noise is constructed over data, which is obtained by running the simulation at strategically placed design points within the trust-region around the current iterate. Then the simulation components of the objective are replaced by their approximations and this analytical and closed-form optimization problem is solved to find the next iterate within the trust-region. Then the trust region is moved and the iterations continue until a satisfactory convergence criteria is met.

KRISHNAMOORTHY, MOHAN↗

DeepHyper: A Python Package for Massively Parallel Hyperparameter Optimization in Machine Learning

Machine learning models are increasingly applied across scientific disciplines, yet their effectiveness often hinges on heuristic decisions—such as data transformations, training strategies, and model architectures—that are not learned by the models themselves. Automating the selection of these heuristics and analyzing their sensitivity is crucial for building robust and efficient learning workflows. DeepHyper addresses this challenge by democratizing hyperparameter optimization, providing accessible tools to streamline and enhance machine learning workflows from a laptop to the largest supercomputer in the world. Building on top of hyperparameter optimization, it unlocks new capabilities around ensembles of models for improved accuracy and uncertainty quantification. All of these organized around efficient parallel computing.

ensemble↗

Quarterly Soil Core and Root Analyses from the Missouri Ozarks AmeriFlux (MOFLUX) Site, Ashland, Missouri, 2017-2023

This dataset contains quarterly soil core measurements from the Missouri Ozarks AmeriFlux (MOFLUX) site located at the University of Missouri’s Thomas H. Baskett Wildlife Research and Education Area near Ashland, Missouri. These data will be used to parameterize an ensemble of MOFLUX-optimized soil carbon-nitrogen models, used to simulate carbon (C) and nitrogen (N) cycling responses to future hydroclimatic scenarios and the trajectory of soil C stocks with concomitant forest decline. Beginning in 2017, eight soil cores were collected approximately quarterly near plot 1 of the southeast transect, near the automated soil respiration flux chambers, from 0–15 cm depth. Data are currently available through 2023 (2017-06-14 to 2023-11-13); additional observations will be appended to this dataset as they become available. Cores were analyzed for gravimetric moisture content, pH, total carbon and nitrogen, texture, microbial biomass carbon and nitrogen, and extractable dissolved organic carbon and nitrogen. This dataset contains one data file in comma separate (*.csv) format. Additional metadata are provided: one data dictionary and a file-level metadata file in comma separate (*.csv) format and a user guide in PDF (*.pdf) format.

54 ENVIRONMENTAL SCIENCES↗

Improving Enzyme Optimum Temperature Prediction with Resampling Strategies and Ensemble Learning

Accurate prediction of the optimal catalytic temperature ( T opt ) of enzymes is vital in biotechnology, as enzymes with high T opt values are desired for enhanced reaction rates. Recently, a machine learning method (temperature optima for microorganisms and enzymes, TOME) for predicting T opt was developed. TOME was trained on a normally distributed data set with a median T opt of 37 °C and less than 5% of T opt values above 85 °C, limiting the method’s predictive capabilities for thermostable enzymes. Due to the distribution of the training data, the mean squared error on T opt values greater than 85 °C is nearly an order of magnitude higher than the error on values between 30 and 50 °C. Here, we apply ensemble learning and resampling strategies that tackle the data imbalance to significantly decrease the error on high T opt values (>85 °C) by 60% and increase the overall R 2 value from 0.527 to 0.632. The revised method, temperature optima for enzymes with resampling (TOMER), and the resampling strategies applied in this work are freely available to other researchers as Python packages on GitHub.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Optimal Power Management for Large-Scale Battery Energy Storage Systems via Bayesian Inference

Large-scale battery energy storage systems (BESS) have found ever-increasing use across industry and society to accelerate clean energy transition and improve energy supply reliability and resilience. However, their optimal power management poses significant challenges: the underlying high-dimensional nonlinear nonconvex optimization lacks computational tractability in real-world implementation, and the uncertainty of the exogenous power demand makes exact optimization difficult. This paper presents a new solution framework to address these bottlenecks. The solution pivots on introducing power-sharing ratios to specify each cell’s power quota from the output power demand. To find the optimal power-sharing ratios, we formulate a nonlinear model predictive control (NMPC) problem to achieve power-loss-minimizing BESS operation while complying with safety, cell balancing, and power supply-demand constraints. We then propose a parameterized control policy for the power-sharing ratios, which utilizes only three parameters, to reduce the computational demand in solving the NMPC problem. This policy parameterization allows us to translate the NMPC problem into a Bayesian inference problem for the sake of 1) computational tractability, and 2) overcoming the nonconvexity of the optimization problem. We leverage the ensemble Kalman inversion technique to solve the parameter estimation problem. Concurrently, a low-level control loop is developed to seamlessly integrate our proposed approach with the BESS to ensure practical implementation. This low-level controller receives the optimal power-sharing ratios, generates output power references for the cells, and maintains a balance between power supply and demand despite uncertainty in output power. We conduct extensive simulations and experiments on a 20-cell prototype to validate the proposed approach.

Battery energy storage systems (BESSs)↗

Machine Learning Approach for Spatiotemporal Multivariate Optimization of Environmental Monitoring Sensor Locations

Abstract Long-term environmental monitoring is critical for managing the soil and groundwater at contaminated sites. Recent improvements in state-of-the-art sensor technology, communication networks, and artificial intelligence have created opportunities to modernize this monitoring activity for automated, fast, robust, and predictive monitoring. In such modernization, it is required that sensor locations be optimized to capture the spatiotemporal dynamics of all monitoring variables as well as to make it cost-effective. The legacy monitoring datasets of the target area are important to perform this optimization. In this study, we have developed a machine-learning approach to optimize sensor locations for soil and groundwater monitoring based on ensemble supervised learning and majority voting. For spatial optimization, Gaussian process regression (GPR) is used for spatial interpolation, while the majority voting is applied to accommodate the multivariate temporal dimension. Results show that the algorithms significantly outperform the random selection of the sensor locations for predictive spatiotemporal interpolation. While the method has been applied to a four-dimensional dataset (with two-dimensional space, time, and multiple contaminants), we anticipate that it can be generalizable to higher-dimensional datasets for environmental monitoring sensor location optimization.

Siddiquee, Masudur R.↗

A catalogue of Locus Algorithm pointings for optimal differential photometry for 23 779 quasars

ABSTRACT This paper presents a catalogue of optimized pointings for differential photometry of 23 779 quasars extracted from the Sloan Digital Sky Survey (SDSS) Catalogue and a Score for each indicating the quality of the Field of View (FoV) associated with that pointing. Observation of millimagnitude variability on a time-scale of minutes typically requires differential observations with reference to an ensemble of reference stars. For optimal performance, these reference stars should have similar colour and magnitude to the target quasar. In addition, the greatest quantity and quality of suitable reference stars may be found by using a telescope pointing which offsets the target object from the centre of the FoV. By comparing each quasar with the stars which appear close to it on the sky in the SDSS Catalogue, an optimum pointing can be calculated, and a figure of merit, referred to as the ‘Score’ is calculated for that pointing. Highly flexible software has been developed to enable this process to be automated and implemented in a distributed computing paradigm, which enables the creation of catalogues of pointings given a set of input targets. Applying this technique to a sample of 40 000 targets from the fourth SDSS quasar catalogue resulted in the production of pointings and Scores for 23 779 quasars based on their magnitudes in the SDSS r-band. This catalogue is a useful resource for observers planning differential photometry studies and surveys of quasars to select those which have many suitable celestial neighbours for differential photometry.

79 ASTRONOMY AND ASTROPHYSICS↗

TurboRVB: A many-body toolkit for ab initio electronic simulations by quantum Monte Carlo

TurboRVB is a computational package for ab initio Quantum Monte Carlo (QMC) simulations of both molecular and bulk electronic systems. The code implements two types of well established QMC algorithms: Variational Monte Carlo (VMC) and diffusion Monte Carlo in its robust and efficient lattice regularized variant. A key feature of the code is the possibility of using strongly correlated many-body wave functions (WFs), capable of describing several materials with very high accuracy, even when standard mean-field approaches [e.g., density functional theory (DFT)] fail. The electronic WF is obtained by applying a Jastrow factor, which takes into account dynamical correlations, to the most general mean-field ground state, written either as an antisymmetrized geminal power with spin-singlet pairing or as a Pfaffian, including both singlet and triplet correlations. This WF can be viewed as an efficient implementation of the so-called resonating valence bond (RVB) Ansatz, first proposed by Pauling and Anderson in quantum chemistry [L. Pauling, The Nature of the Chemical Bond (Cornell University Press, 1960)] and condensed matter physics [P.W. Anderson, Mat. Res. Bull 8, 153 (1973)], respectively. The RVB Ansatz implemented in TurboRVB has a large variational freedom, including the Jastrow correlated Slater determinant as its simplest, but nontrivial case. Moreover, it has the remarkable advantage of remaining with an affordable computational cost, proportional to the one spent for the evaluation of a single Slater determinant. Therefore, its application to large systems is computationally feasible. The WF is expanded in a localized basis set. Several basis set functions are implemented, such as Gaussian, Slater, and mixed types, with no restriction on the choice of their contraction. The code implements the adjoint algorithmic differentiation that enables a very efficient evaluation of energy derivatives, comprising the ionic forces. Thus, one can perform structural optimizations and molecular dynamics in the canonical NVT ensemble at the VMC level. For the electronic part, a full WF optimization (Jastrow and antisymmetric parts together) is made possible, thanks to state-of-the-art stochastic algorithms for energy minimization. In the optimization procedure, the first guess can be obtained at the mean-field level by a built-in DFT driver. The code was efficiently parallelized by using a hybrid MPI-OpenMP protocol, which is also an ideal environment for exploiting the computational power of modern Graphics Processing Unit accelerators.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine-learning accelerated geometry optimization in molecular simulation

Geometry optimization is an important part of both computational materials and surface science because it is the path to finding ground state atomic structures and reaction pathways. These properties are used in the estimation of thermodynamic and kinetic properties of molecular and crystal structures. This process is slow at the quantum level of theory because it involves an iterative calculation of forces using quantum chemical codes such as density functional theory (DFT), which are computationally expensive and which limit the speed of the optimization algorithms. It would be highly advantageous to accelerate this process because then one could do either the same amount of work in less time or more work in the same time. Here, we provide a neural network (NN) ensemble based active learning method to accelerate the local geometry optimization for multiple configurations simultaneously. We illustrate the acceleration on several case studies including bare metal surfaces, surfaces with adsorbates, and nudged elastic band for two reactions. In all cases, the accelerated method requires fewer DFT calculations than the standard method. In addition, we provide an Atomic Simulation Environment (ASE)-optimizer Python package to make the usage of the NN ensemble active learning for geometry optimization easier.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Mechanistic Insights and Rational Design of Ca-Doped CeO 2 Catalyst for Acetic Acid Ketonization

Carboxylic acid ketonization has recently gained significant attention to produce biomass-derived hydrocarbon fuels as it not only removes the highly reactive carboxylic functional group but also increases the size of the carbon chain. In this work, Ca-doped CeO 2 -based catalysts were investigated for acetic acid ketonization using a combined experimental and computational approach. Acetic acid conversion was performed across a range of temperatures including higher temperatures relevant to catalytic hot gas filtration (450 °C). Ca addition slightly decreases overall acetic acid ketonization reactivity yet stabilizes the catalyst at the higher temperatures necessary for catalytic hot gas filtration. From density functional theory calculations of the ketonization reaction mechanism, the C–C coupling and water formation steps are identified as two of the most energy-consuming steps on a CeO 2 surface with a proximal oxygen vacancy and the presence of a Ca dopant stabilizes the key intermediates. Calculations predict an optimal structure comprising three Ca ensembles to minimize the reaction free energies for C–C coupling and water formation steps. These findings provide a priori information to guide future experiments for ketonization catalyst design and development.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Prediction of Distributed River Sediment Respiration Rates Using Community-Generated Data and Machine Learning

River sediment microbial respiration is a key indicator of ecosystem functioning and the biogeochemical fluxes across this critical zone link surface and subsurface waters. As such, there is tremendous interest in measuring and mapping these respiration rates. Respiration observations are expensive and labor intensive; there is limited data available to the community. An open science, collaborative initiative is collecting samples for respiration rate analysis and multi-scale metadata; this evolving data set is being used for making machine learning (ML) predictions at unsampled sites to help inform continued community engagement. However, it is a challenge to find an optimum configuration for ML models to work with this feature-rich (i.e., 100+ possible input variables) data set. Here, we present results from a two-tiered approach to managing the analysis of this complex data set: (a) a stacked ensemble of models that automatically optimizes hyperparameters and manages the training of many models and (b) feature permutation importance to detect the most important features in the models. The major elements of this workflow are modular, portable, open, and cloud-based thus making this implementation a potential template for other applications. The models developed here predict that sediment organic matter chemistry is one of the most important features for predicting sediment respiration rate. Other larger-scale, important features fall into the categories of climatic, ecological, geological, and fluvial settings. Leveraging these larger-scale features to generate data-driven estimates of river sediment respiration rates reveals spatially consistent but heterogeneous patterns across the river network of the Columbia River Basin.

54 ENVIRONMENTAL SCIENCES↗

Validating first-principles molecular dynamics calculations of oxide/water interfaces with x-ray reflectivity data

Metal oxide/water interfaces play a crucial role in many electrochemical and photocatalytic processes, such as photoelectrochemical water splitting, the creation of fuel from sunlight, and electrochemical CO 2 reduction. First-principles electronic structure calculations can reveal unique insights into these processes, such as the role of the alignment of the oxide electronic energy levels with those of liquid water. An essential prerequisite for the success of such calculations is the ability to predict accurate structural models of these interfaces, which in turn requires careful experimental validation. Here we report a general, quantitative validation protocol for first-principles molecular dynamics simulations of oxide/aqueous interfaces. The approach makes direct comparisons of interfacial x-ray reflectivity (XR) signals from experimental measurements and those obtained from ab initio simulations with semilocal and van der Waals functionals. The protocol is demonstrated here for the case of the Al 2 O 3 (001)/water interface, one of the simplest oxide/water interfaces. We discuss the technical requirements needed for validation, including the choice of the density functional, the simulation cell size, and the optimal choice of the thermodynamic ensemble. Our results establish a general paradigm for the validation of structural models and interactions at solid/water interfaces derived from first-principles simulations. Furthermore, while there is qualitative agreement between the simulated structures and the experimental best-fit structure, direct comparisons of simulated and measured XR intensities show quantitative discrepancies that derive from both bulk regions (i.e., alumina and water) as well as the interfacial region, highlighting the need for accurate density functionals to properly describe interfacial interactions. Our results show that XR data are sensitive not only to the atomic structure (i.e., the atom locations) but also to the electron-density distributions in both the substrate and at the interface.

36 MATERIALS SCIENCE↗

Prediction of the development of islet autoantibodies through integration of environmental, genetic, and metabolic markers

The Environmental Determinants of the Diabetes in the Young (TEDDY) study has prospectively followed, from birth, children at increased genetic risk of type 1 diabetes. We evaluated the potential of machine learning to identify new biomarkers that predict imminent (within 6 months) development of persistent islet autoantibodies to insulin, GAD or IA-2 in TEDDY participants through integration of time-invariant risk factors with time-varying metabolomics. The predictive modeling was initiated with over 220 potential biomarkers; through ensemble-based feature evaluation, the optimal model included 42 biomarkers, returning a cross-validated receiver operating characteristic area under the curve of 0.74. The model identified a principal set of 20 time-invariant markers, including 16 single nucleotide polymorphisms and two HLA-DR genotypes, gestational age, and exposure to a prebiotic formula. Integration of the metabolome identified 22 high-priority metabolites and lipids, including adipic acid and ceramide d42:0, that predicted development of islet autoantibodies, dependent upon the time horizon. The majority (86%) of metabolites that predicted development of islet autoantibodies belonged to 3 pathways: lipid oxidation, phospholipase A2 signaling, and pentose phosphate pathway. TEDDY data suggest that these metabolic processes may play a role in triggering islet autoimmunity.

59 BASIC BIOLOGICAL SCIENCES↗

Bioactive scaffolds with enhanced supramolecular motion promote recovery from spinal cord injury

The signaling of cells by scaffolds of synthetic molecules that mimic proteins is known to be effective in the regeneration of tissues. Here, we describe peptide amphiphile supramolecular polymers containing two distinct signals and test them in a mouse model of severe spinal cord injury. One signal activates the transmembrane receptor β1-integrin and a second one activates the basic fibroblast growth factor 2 receptor. By mutating the peptide sequence of the amphiphilic monomers in nonbioactive domains, we intensified the motions of molecules within scaffold fibrils. This resulted in notable differences in vascular growth, axonal regeneration, myelination, survival of motor neurons, reduced gliosis, and functional recovery. Here, we hypothesize that the signaling of cells by ensembles of molecules could be optimized by tuning their internal motions.

36 MATERIALS SCIENCE↗

Systems and methods for modeling water quality

A system, method, device and computer-readable medium for creating an ensemble model of water quality. The ensemble model is generated by determining a set of optimal component models for spectral regions of a body of water, and combining the optimal models. The optimal models can be based on remote sensing data, including satellite imagery. A K-fold partition approach or a global approach can be used to determine the optimal component models, and the optimal component models can be combined through spectral space partition rules to generate an ensemble model of water quality. The ensemble model not only has improved water quality prediction ability, but also has strong spatial and temporal extensibility. The spatial and temporal extensibility of the ensemble model is fundamentally important and desirable for long-term and large-scale remote sensing monitoring and assessment of water quality.

Liu, Hongxing↗

Scalable deep learning for watershed model calibration

Watershed models such as the Soil and Water Assessment Tool (SWAT) consist of high-dimensional physical and empirical parameters. These parameters often need to be estimated/calibrated through inverse modeling to produce reliable predictions on hydrological fluxes and states. Existing parameter estimation methods can be time consuming, inefficient, and computationally expensive for high-dimensional problems. In this paper, we present an accurate and robust method to calibrate the SWAT model (i.e., 20 parameters) using scalable deep learning (DL). We developed inverse models based on convolutional neural networks (CNN) to assimilate observed streamflow data and estimate the SWAT model parameters. Scalable hyperparameter tuning is performed using high-performance computing resources to identify the top 50 optimal neural network architectures. We used ensemble SWAT simulations to train, validate, and test the CNN models. We estimated the parameters of the SWAT model using observed streamflow data and assessed the impact of measurement errors on SWAT model calibration. We tested and validated the proposed scalable DL methodology on the American River Watershed, located in the Pacific Northwest-based Yakima River basin. Our results show that the CNN-based calibration is better than two popular parameter estimation methods (i.e., the generalized likelihood uncertainty estimation [GLUE] and the dynamically dimensioned search [DDS], which is a global optimization algorithm). For the set of parameters that are sensitive to the observations, our proposed method yields narrower ranges than the GLUE method but broader ranges than values produced using the DDS method within the sampling range even under high relative observational errors. The SWAT model calibration performance using the CNNs, GLUE, and DDS methods are compared using R 2 and a set of efficiency metrics, including Nash-Sutcliffe, logarithmic Nash-Sutcliffe, Kling-Gupta, modified Kling-Gupta, and non-parametric Kling-Gupta scores, computed on the observed and simulated watershed responses. The best CNN-based calibrated set has scores of 0.71, 0.75, 0.85, 0.85, 0.86, and 0.91. The best DDS-based calibrated set has scores of 0.62, 0.69, 0.8, 0.77, 0.79, and 0.82. The best GLUE-based calibrated set has scores of 0.56, 0.58, 0.71, 0.7, 0.71, and 0.8. The scores above show that the CNN-based calibration leads to more accurate low and high streamflow predictions than the GLUE and DDS sets. Our research demonstrates that the proposed method has high potential to improve our current practice in calibrating large-scale integrated hydrologic models.

54 ENVIRONMENTAL SCIENCES↗

A novel machine learning-based optimization algorithm (ActivO) for accelerating simulation-driven engine design

A novel design optimization approach (ActivO) that employs an ensemble of machine learning algorithms is presented. The proposed approach is a surrogate-based scheme, where the predictions of a weak leaner and a strong learner are utilized within an active learning loop. The weak learner is used to identify promising regions within the design space to explore, while the strong learner is used to determine the exact location of the optimum within promising regions. For each design iteration, exploration is done by randomly selecting evaluation points within regions where the weak learner-predicted fitness is high. The global optimum obtained by using the strong learner as a surrogate is also evaluated to enable rapid convergence once the most promising region has been identified. First, the performance of ActivO was compared against five other optimizers on a cosine mixture function with 25 local optima and one global optimum. In the second problem, the objective was to minimize indicated specific fuel consumption of a compression-ignition internal combustion (IC) engine while adhering to desired constraints associated with in-cylinder pressure and emissions. In this work, the efficacy of the proposed approach is compared to that of a genetic algorithm, which is widely used within the internal combustion engine community for engine optimization, showing that ActivO reduces the number of function evaluations needed to reach the global optimum, and thereby time-to-design by 80%. Furthermore, the optimization of engine design parameters leads to savings of around 1.9% in energy consumption, while maintaining operability and acceptable pollutant emissions.

97 MATHEMATICS AND COMPUTING↗

CeO2 Nanoparticle Doping as a Probe of Active Site Speciation in the Catalytic Hydrolysis of Organophosphates

Organophosphate hydrolysis is important for degrading environmentally harmful compounds and recovering phosphate ions in biological molecules. CeO2 nanocrystals have been well-studied for dephosphorylation via hydrolysis owing to the accessible and tunable distribution of Ce3+ and Ce4+ ions. However, there remains uncertainty in the literature regarding which surface defect properties direct catalytic activity, such as the Ce3+/Ce4+ distribution, oxygen vacancies, faceting, and dopants, and to what degree they contribute to efficient hydrolysis. Trivalent (M3+) dopants serve as a tool for manipulating defects, including the concentration of Ce3+ and oxygen vacancies, thereby influencing the hydrolytic activity of CeO2. Herein, trivalent metal ions (M = Y3+, Cr3+, In3+, and Gd3+) were employed to modulate the active sites on the CeO2 nanocrystal surface, and the effects of each metal dopant on the cerium oxide active sites for organophosphate hydrolysis were investigated. M-doped CeO2 nanoparticles were synthesized via hydrothermal methods, followed by annealing to remove ligands and prime the nanocrystal surface for catalysis. Catalytic performance was evaluated using dimethyl-p¬-nitrophenyl phosphate (DMNP) as a model organophosphate substrate, with degradation monitored over time using UV-visible absorption spectroscopy. Powder X-ray diffraction (PXRD), X-ray photoelectron spectroscopy (XPS), and Raman spectroscopy revealed successful doping of CeO2 in all cases, albeit with distinctive characteristics demonstrating how M3+ dopants affect catalysis. We show that CeO2 exhibits high sensitivity to dopants that generate lattice strain, Ce3+ ions, and oxygen vacancy defects. Consequently, achieving high catalytic efficiency within CeO2 requires a balanced active site ensemble, wherein defects are maintained at optimal concentrations and distributions on the nanocrystal surface.

Miura-Stempel, Emily L.↗