Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “ensemble methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Scale setting the Möbius domain wall fermion on gradient-flowed HISQ action using the omega baryon mass and the gradient-flow scales 𝑡 0 and 𝑤 0

We report on a subpercent scale determination using the omega baryon mass and gradient-flow methods. The calculations are performed on 22 ensembles of 𝑁 𝑓 =2 +1 +1 highly improved, rooted staggered sea-quark configurations generated by the MILC and CalLat Collaborations. The valence quark action used is Möbius domain wall fermions solved on these configurations after a gradient-flow smearing is applied with a flowtime of 𝑡 gf = 1 in lattice units. The ensembles span four lattice spacings in the range 0.06 ≲ 𝑎 ≲0.15 fm, six pion masses in the range 130 ≲ 𝑚 𝜋 ≲ 400 MeV and multiple lattice volumes. On each ensemble, the gradient-flow scales 𝑡 0 /𝑎 2 and 𝑤 0 /𝑎 and the omega baryon mass 𝑎⁢𝑚 Ω are computed. The dimensionless product of these quantities is then extrapolated to the continuum and infinite volume limits and interpolated to the physical light, strange and charm quark mass point in the isospin limit, resulting in the determination of $\sqrt{t}_0$ = 0.1422⁢(14) fm and 𝑤 0 = 0.1709⁢(11) fm with all sources of statistical and systematic uncertainty accounted for. The dominant uncertainty in both results is the stochastic uncertainty, though for $\sqrt{t}_0$ there are comparable continuum extrapolation uncertainties. For 𝑤 0 , there is a clear path for a few-per-mille uncertainty just through improved stochastic precision, as recently obtained by the Budapest-Marseille-Wuppertal Collaboration.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Contact in the Unitary Fermi Gas across the Superfluid Phase Transition

A quantity known as the contact is a fundamental thermodynamic property of quantum many-body systems with short-range interactions. Determination of the temperature dependence of the contact for the unitary Fermi gas of infinite scattering length has been a major challenge, with different calculations yielding qualitatively different results. Here we use finite-temperature auxiliary-field quantum Monte Carlo (AFMC) methods on the lattice within the canonical ensemble to calculate the temperature dependence of the contact for the homogeneous spin-balanced unitary Fermi gas. We extrapolate to the continuum limit for 40, 66, and 114 particles, eliminating systematic errors due to finite-range effects. We observe a dramatic decrease in the contact as the superfluid critical temperature is approached from below, followed by a gradual weak decrease as the temperature increases in the normal phase. Our theoretical results are in excellent agreement with the most recent precision ultracold atomic gas experiments. Here, we also present results for the energy as a function of temperature in the continuum limit.

74 ATOMIC AND MOLECULAR PHYSICS↗

Structure–activity relationship-based chemical classification of highly imbalanced Tox21 datasets

Abstract The specificity of toxicant-target biomolecule interactions lends to the very imbalanced nature of many toxicity datasets, causing poor performance in Structure–Activity Relationship (SAR)-based chemical classification. Undersampling and oversampling are representative techniques for handling such an imbalance challenge. However, removing inactive chemical compound instances from the majority class using an undersampling technique can result in information loss, whereas increasing active toxicant instances in the minority class by interpolation tends to introduce artificial minority instances that often cross into the majority class space, giving rise to class overlapping and a higher false prediction rate. In this study, in order to improve the prediction accuracy of imbalanced learning, we employed SMOTEENN, a combination of Synthetic Minority Over-sampling Technique (SMOTE) and Edited Nearest Neighbor (ENN) algorithms, to oversample the minority class by creating synthetic samples, followed by cleaning the mislabeled instances. We chose the highly imbalanced Tox21 dataset, which consisted of 12 in vitro bioassays for > 10,000 chemicals that were distributed unevenly between binary classes. With Random Forest (RF) as the base classifier and bagging as the ensemble strategy, we applied four hybrid learning methods, i.e., RF without imbalance handling (RF), RF with Random Undersampling (RUS), RF with SMOTE (SMO), and RF with SMOTEENN (SMN). The performance of the four learning methods was compared using nine evaluation metrics, among which F 1 score, Matthews correlation coefficient and Brier score provided a more consistent assessment of the overall performance across the 12 datasets. The Friedman’s aligned ranks test and the subsequent Bergmann-Hommel post hoc test showed that SMN significantly outperformed the other three methods. We also found that a strong negative correlation existed between the prediction accuracy and the imbalance ratio (IR), which is defined as the number of inactive compounds divided by the number of active compounds. SMN became less effective when IR exceeded a certain threshold (e.g., > 28). The ability to separate the few active compounds from the vast amounts of inactive ones is of great importance in computational toxicology. This work demonstrates that the performance of SAR-based, imbalanced chemical toxicity classification can be significantly improved through the use of data rebalancing.

Idakwo, Gabriel↗

Daily evapotranspiration changes during heatwaves at 32 NEON sites, 2019-2021

This dataset provides partitioned evapotranspiration (ET, the combined loss of water from soil and plant surfaces) anomalies during heatwave events—soil evaporation (E) and transpiration (T)—for 268 heatwave events across 32 National Ecological Observatory Network (NEON) flux sites in the contiguous United States from 2019–2021. Using an ensemble of four high-frequency turbulence methods (Flux-variance Similarity, Conditional Eddy Covariance [CEC], CEC with Water-Use Efficiency, and Conditional Eddy Accumulation; see Zahn and Bou-Zeid 2024), half-hourly transpiration-to-evapotranspiration (T/ET) ratios were derived from 20 hertz (Hz, cycles per second) eddy covariance measurements of carbon dioxide (CO₂) and water vapor (H₂O) concentrations. The dataset spans six vegetation types including evergreen and deciduous forests, grasslands, cultivated crops, shrublands, and emergent herbaceous wetlands. Data Package Contents: The dataset includes a single CSV (comma-separated values) file containing daily anomalies (deviations from baseline conditions) for transpiration (Delta_T), evaporation (Delta_E), total evapotranspiration (Delta_ET), and T/ET ratio (Delta_T_ET) during each day of identified heatwave events. The file also includes site codes, dates, heatwave event identifiers, and day-of-heatwave indicators. The CSV file can be opened with spreadsheet software (Microsoft Excel, Google Sheets) or programming environments (Python, R, MATLAB). This resource enables researchers to investigate ecosystem-specific responses to thermal extremes, validate land surface model partitioning of ET fluxes, and examine feedbacks between water cycling and surface energy balance during heatwaves. The dataset is particularly valuable for studies linking vegetation hydraulic strategies to climate resilience, as it captures the divergent responses of shallow-rooted versus deep-rooted ecosystems. Potential applications include improving drought early warning systems, informing irrigation management strategies, and advancing our mechanistic understanding of land-atmosphere interactions under extreme heat conditions.

Day of Heatwave↗

Maximum Entropy Theory of Multiscale Coarse-Graining via Matching Thermodynamic Forces: Application to a Molecular Crystal (TATB)

The MSCG/FM (multiscale coarse-graining via force-matching) approach is an efficient supervised machine learning method to develop microscopically informed coarse-grained (CG) models. Here we present a theory based on the principle of maximum entropy (PME) enveloping the existing MSCG/FM approaches. This theory views the MSCG/FM method as a special case of matching the thermodynamic forces from the extended ensemble described by the set of thermodynamic (relevant) system coordinates. This set may include CG coordinates, the stress tensor, applied external fields, and so forth, and may be characterized by nonequilibrium conditions. Following the presentation of the theory, we discuss the consistent matching of both bonded and nonbonded interactions. The proposed PME formulation is used as a starting point to extend the MSCG/FM method to the constant strain ensemble, which together with the explicit matching of the bonded forces is better suited for coarse-graining anisotropic media at a submolecular resolution. The theory is demonstrated by performing the fine coarse-graining of crystalline 1,3,5-triamino-2,4,6-trinitrobenzene (TATB), a well-known insensitive molecular energetic material, which exhibits highly anisotropic mechanical properties.

1,3,5-triamino-2,4,6-trinitrobenzene↗

Structural Rearrangements of Subnanometer Cu Oxide Clusters Govern Catalytic Oxidation

Sub-nanometer metal oxide clusters are very important materials that are widely used, for example, in catalysis or electronic devices such as sensors. Hence, it is critical to understand the atomic structures and properties of sub-nanometer metal oxide clusters under a reactive gas environment, such as O 2 . We consider here experimentally accessible precise-size Cu clusters (Cu 4 ) supported on partially hydroxylated amorphous alumina and show that such clusters can access, in catalytic conditions at high temperature under a pressure of O 2 , a large ensemble of oxidized structures, representing a large variety of oxygen content, of geometries in link with the support, and of catalytic activities for oxidation reactions, as seen from their reducibilities. A grand canonical basin hopping method based on first-principles energy reveals an ensemble of 24 configurations for the Cu 4 O x cluster of low free energy, less than 0.8 eV above the global minimum. The low free energy ensemble consists of clusters of different stoichiometries, which are mainly Cu 4 O 3 and Cu 4 O 4 in the temperature range of 200–400 °C and under a pressure of 0.5 bar of O 2 . The presence of several competitive isomers at each composition implies that cluster fluxionality impacts the phase diagram, which should be ensemble-averaged. In terms of catalytic oxidation activity, Cu 4 O 3 isomers present highly variable O abstraction energy: the most stable isomers are inactive for alkane oxidative dehydrogenation, but isomerization to metastable isomers, that proceed with low barrier, enable to create active configurations with low O abstraction energy. O atoms with the lowest anionic character, and thus of more electrophilic nature, present the best oxidation capability. In contrast, all Cu 4 O 4 isomers show a low O abstraction energy and a high potential catalytic activity. Here, this manuscript demonstrates the unique structural and electronic properties of sub-nano Cu oxide clusters and illustrates the critical roles of configuration ensembles and rearrangement to highly reactive metastable cluster isomers in nanocatalysis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Adaptive Uncertainty Quantification for Stochastic Hyperbolic Conservation Laws

Here, we propose a predictor-corrector adaptive method for the study of hyperbolic partial differential equations (PDEs) under uncertainty. Constructed around the framework of stochastic finite volume (SFV) methods, our approach circumvents sampling schemes or simulation ensembles while also preserving fundamental properties, in particular hyperbolicity of the resulting systems and conservation of the discrete solutions. Furthermore, we augment the existing SFV theory with a priori convergence results for statistical quantities, in particular push-forward densities, which we demonstrate through numerical experiments. By linking refinement indicators to regions of the physical and stochastic spaces, we drive anisotropic refinements of the discretizations, introducing new degrees of freedom where deemed profitable. To illustrate our proposed method, we consider a series of numerical examples for nonlinear hyperbolic PDEs based on Burgers’ and Euler’s equations.

97 MATHEMATICS AND COMPUTING↗

Protein Conformational States—A First Principles Bayesian Method

Automated identification of protein conformational states from simulation of an ensemble of structures is a hard problem because it requires teaching a computer to recognize shapes. We adapt the naïve Bayes classifier from the machine learning community for use on atom-to-atom pairwise contacts. The result is an unsupervised learning algorithm that samples a ‘distribution’ over potential classification schemes. We apply the classifier to a series of test structures and one real protein, showing that it identifies the conformational transition with >95% accuracy in most cases. A nontrivial feature of our adaptation is a new connection to information entropy that allows us to vary the level of structural detail without spoiling the categorization. This is confirmed by comparing results as the number of atoms and time-samples are varied over 1.5 orders of magnitude. Further, the method’s derivation from Bayesian analysis on the set of inter-atomic contacts makes it easy to understand and extend to more complex cases.

97 MATHEMATICS AND COMPUTING↗

Dressed-State Hamiltonian Engineering in a Strongly Interacting Solid-State Spin Ensemble

In quantum science applications, ranging from many-body physics to quantum metrology, dipolar interactions in spin ensembles are often controlled via Floquet engineering. However, this technique typically reduces the interaction strength between spins and effectively weakens the coupling to a target sensing field, limiting the metrological sensitivity. In this Letter, we develop and demonstrate an alternative method that directly tunes the native dipolar interaction in an ensemble of nitrogen-vacancy (NV) centers in diamond, thereby overcoming these limitations inherent to Floquet engineering. Our approach utilizes dressed-state qubit encoding under a bias magnetic field applied perpendicular to the crystal lattice orientation. This method leads to a 3.2× enhancement of the dimensionless coherence parameter JT 2 compared to state-of-the-art Floquet engineering and a 2.6× (8.3 dB) enhanced sensitivity in ac magnetometry. Furthermore, our results provide a powerful Hamiltonian engineering tool for future studies with NV ensembles and other interacting higher-spin (S > $\frac{1}{2}$) systems.

Quantum control↗

Feasibility of DEIM for retrieving the initial field via dimensionality reduction

When parameter estimation is solved in a high-dimensional space, the dimensionality reduction strategy becomes the primary consideration for alleviating the tremendous computational cost. Here, the discrete empirical interpolation method (DEIM) is explored to retrieve the initial condition (IC) by combining the polynomial chaos (PC) based ensemble Kalman filter (i.e. PC-EnKF), where a non-intrusive PC expansion is considered as a surrogate model in place of the forward model in the prediction step of the ensemble Kalman filter, resulting in fewer forward model integrations but with a comparable accuracy as Monte Carlo-based approaches. The DEIM acts as a hyper-reduction tool to provide the low-dimensional input for the high-dimensional initial field, which can be reconstructed using the information on the sparse interpolation grid points that is adaptively obtained through PC-EnKF data assimilation method. Thus an innovative framework to reconstruct the IC is developed. The detailed procedure at each assimilation iteration includes: the determination of the spatial interpolation points, the estimation of the initial values on the interpolation locations using the optimal observations, and the reconstruction of IC in the full space. The current study uses the reconstruction field of initial conditions of the Navier-Stokes equations as an example to illustrate the efficacy of our method. The experimental results demonstrate the proposed algorithm achieves a satisfactory reconstruction for the initial field. The proposed method helps to extend the applicable area of DEIM in solving inverse problems.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

On the Rapid Calculation of Binding Affinities for Antigen and Antibody Design and Affinity Maturation Simulations

The accurate and efficient calculation of protein-protein binding affinities is an essential component in antibody and antigen design and optimization, and in computer modeling of antibody affinity maturation. Such calculations remain challenging despite advances in computer hardware and algorithms, primarily because proteins are flexible molecules, and thus, require explicit or implicit incorporation of multiple conformational states into the computational procedure. The astronomical size of the amino acid sequence space further compounds the challenge by requiring predictions to be computed within a short time so that many sequence variants can be tested. In this study, we compare three classes of methods for antibody/antigen (Ab/Ag) binding affinity calculations: (i) a method that relies on the physical separation of the Ab/Ag complex in equilibrium molecular dynamics (MD) simulations, (ii) a collection of 18 scoring functions that act on an ensemble of structures created using homology modeling software, and (iii) methods based on the molecular mechanics-generalized Born surface area (MM-GBSA) energy decomposition, in which the individual contributions of the energy terms are scaled to optimize agreement with the experiment. When applied to a set of 49 antibody mutations in two Ab/HIV gp120 complexes, all of the methods are found to have modest accuracy, with the highest Pearson correlations reaching about 0.6. In particular, the most computationally intensive method, i.e., MD simulation, did not outperform several scoring functions. The optimized energy decomposition methods provided marginally higher accuracy, but at the expense of requiring experimental data for parametrization. Within each method class, we examined the effect of the number of independent computational replicates, i.e., modeled structures or reinitialized MD simulations, on the prediction accuracy. We suggest using about ten modeled structures for scoring methods, and about five simulation replicates for MD simulations as a rule of thumb for obtaining reasonable convergence. We anticipate that our study will be a useful resource for practitioners working to incorporate binding affinity calculations within their protein design and optimization process.

59 BASIC BIOLOGICAL SCIENCES↗

Pooling Data Improves Multimodel IDF Estimates over Median-Based IDF Estimates: Analysis over the Susquehanna and Florida

Traditional multimodel methods for estimating future changes in precipitation intensity, duration, and frequency (IDF) curves rely on mean or median of models’ IDF estimates. Such multimodel estimates are impaired by large estimation uncertainty, shadowing their efficacy in planning efforts. Here, assuming that each climate model is one representation of the underlying data generating process, i.e., the Earth system, we propose a novel extension of current methods through pooling model data: (i) evaluate performance of climate models in simulating the spatial and temporal variability of the observed annual maximum precipitation (AMP), (ii) bias-correct and pool historical and future AMP data of reasonably performing models, and (iii) compute IDF estimates in a nonstationary framework from pooled historical and future model data. Pooling enhances fitting of the extreme value distribution to the data and assumes that data from reasonably performing models represent samples from the “true” underlying data generating distribution. Through Monte Carlo simulations with synthetic data, we show that return periods derived from pooled data have smaller biases and lesser uncertainty than those derived from ensembles of individual model data. We apply this method to NA-CORDEX models to estimate changes in 24-h precipitation intensity–frequency (PIF) estimates over the Susquehanna watershed and Florida peninsula. Our approach identifies significant future changes at more stations compared to median-based PIF estimates. The analysis suggests that almost all stations over the Susquehanna and at least two-thirds of the stations over the Florida peninsula will observe significant increases in 24-h precipitation for 2–100-yr return periods.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The Power of Many: An Ensemble Approach to Spectral Similarity

Quantifying the similarity between two mass spectra─a known reference mass spectrum and an unidentified sample mass spectrum─is at the heart of compound identification workflows in gas chromatography–mass spectrometry (GC-MS). The reference spectrum most like the sample is assigned as its identification (provided some quantitative similarity threshold is met, e.g., 80%) and thus accurately measuring similarity is essential. Significant research has gone toward developing metrics for this purpose, each of which has attempted to improve upon existing methods by incorporating GC-MS-specific information (e.g., peak ratios or retention times) or adopting various statistical and algorithmic frameworks. While this active development has led to a plethora of similarity metrics with demonstrated value across different contexts, the unfortunate consequence has been confusion surrounding which metric should be used as a global standard. No such metric is currently accepted as the standard method because different metrics have demonstrated optimal performance in different contexts. In this work, we propose an ensemble approach to spectral similarity scoring that combines the collective information from across existing similarity metrics to form an improved, globally representative similarity metric as a step toward establishing a global standard method. In conclusion, the resulting ensemble metrics are evaluated on over 88,000 spectra of varying complexity and demonstrate improved abilities to accurately rank the correct reference spectrum as the top-matching candidate for a sample relative to the rankings generated by individual similarity scores.

Carbohydrates↗

Conformalized-KANs: Uncertainty Quantification with Coverage Guarantees for Kolmogorov-Arnold Networks (KANs) in Scientific Machine Learning

This paper explores uncertainty quantification (UQ) methods in the context of Kolmogorov–Arnold Networks (KANs). We apply an ensemble approach to KANs to obtain a heuristic measure of UQ, enhancing interpretability and robustness in modeling complex functions. Building on this, we introduce Conformalized-KANs, which integrate conformal prediction, a distribution-free UQ technique, with KAN ensembles to generate calibrated prediction intervals with guaranteed coverage.} Extensive numerical experiments are conducted to evaluate the effectiveness of these methods, focusing particularly on the robustness and accuracy of the prediction intervals under various hyperparameter settings. We show that the conformal KAN predictions can be applied to recent extensions of KANs, including Finite Basis KANs (FBKANs) and multifideilty KANs (MFKANs). The results demonstrate the potential of our approaches to significantly improve the reliability and applicability of KANs in scientific machine learning.

• Artificial intelligence (AI) / machine learning ↗

Estimating Subhourly Inverter Clipping Loss From Satellite-Derived Irradiance Data

Photovoltaic system production simulations are conventionally run using hourly weather datasets. Hourly simulations are sufficiently accurate to predict the majority of long-term system behavior but cannot resolve high-frequency effects like inverter clipping caused by short-duration irradiance variability. Direct modeling of this subhourly clipping error is only possible for the few locations with high-resolution irradiance datasets. This paper describes a method of predicting the magnitude of this error using a machine learning regressor ensemble model, comprised of a random forest and an XGBoost model, and 30-minute satellite irradiance data. The method predicts a correction for each 30-minute interval with the potential to roll up into 60-minute corrections to match an hourly energy model. The model is trained and validated at locations where the error can be directly simulated from 1-minute ground data. The validation shows low bias at most ground station locations. The model is also applied to gridded satellite irradiance to produce a heatmap of the estimated clipping error across the United States. Finally, the relative importance of each predictor satellite variable is retrieved from the model and discussed.

41 EE - Solar Energy Technologies Office (EE-4S)↗

Revealing the Statistics of Extreme Events Hidden in Short Weather Forecast Data

Extreme weather events have significant consequences, dominating the impact of climate on society. While high-resolution weather models can forecast many types of extreme events on synoptic timescales, long-term climatological risk assessment is an altogether different problem. A once-in-a-century event takes, on average, 100 years of simulation time to appear just once, far beyond the typical integration length of a weather forecast model. Therefore, this task is left to cheaper, but less accurate, low-resolution or statistical models. But there is untapped potential in weather model output: despite being short in duration, weather forecast ensembles are produced multiple times a week. Integrations are launched with independent perturbations, causing them to spread apart over time and broadly sample phase space. Collectively, these integrations add up to thousands of years of data. We establish methods to extract climatological information from these short weather simulations. Using ensemble hindcasts by the European Center for Medium-range Weather Forecasting archived in the subseasonal-to-seasonal (S2S) database, we characterize sudden stratospheric warming (SSW) events with multi-centennial return times. Consistent results are found between alternative methods, including basic counting strategies and Markov state modeling. By carefully combining trajectories together, we obtain estimates of SSW frequencies and their seasonal distributions that are consistent with reanalysis-derived estimates for moderately rare events, but with much tighter uncertainty bounds, and which can be extended to events of unprecedented severity that have not yet been observed historically. These methods hold potential for assessing extreme events throughout the climate system, beyond this example of stratospheric extremes.

58 GEOSCIENCES↗

Machine Learning Emulation of Spatial Deposition from a Multi-Physics Ensemble of Weather and Atmospheric Transport Models

In the event of an accidental or intentional hazardous material release in the atmosphere, researchers often run physics-based atmospheric transport and dispersion models to predict the extent and variation of the contaminant spread. These predictions are imperfect due to propagated uncertainty from atmospheric model physics (or parameterizations) and weather data initial conditions. Ensembles of simulations can be used to estimate uncertainty, but running large ensembles is often very time consuming and resource intensive, even using large supercomputers. In this paper, we present a machine-learning-based method which can be used to quickly emulate spatial deposition patterns from a multi-physics ensemble of dispersion simulations. We use a hybrid linear and logistic regression method that can predict deposition in more than 100,000 grid cells with as few as fifty training examples. Logistic regression provides probabilistic predictions of the presence or absence of hazardous materials, while linear regression predicts the quantity of hazardous materials. The coefficients of the linear regressions also open avenues of exploration regarding interpretability—the presented model can be used to find which physics schemes are most important over different spatial areas. A single regression prediction is on the order of 10,000 times faster than running a weather and dispersion simulation. However, considering the number of weather and dispersion simulations needed to train the regressions, the speed-up achieved when considering the whole ensemble is about 24 times. Ultimately, this work will allow atmospheric researchers to produce potential contamination scenarios with uncertainty estimates faster than previously possible, aiding public servants and first responders.

97 MATHEMATICS AND COMPUTING↗

CONCURRENT, CONDENSED STEIN VARIATIONAL GRADIENT DESCENT FOR UNCERTAINTY QUANTIFICATION OF NEURAL NETWORKS

In this work, we propose a Stein variational gradient descent (SVGD) method to concurrently sparsify, train, and provide uncertainty quantification (UQ) of a complexly parameterized model, such as a neural network (NN). It employs a graph reconciliation and condensation process to reduce complexity and increase similarity in the Stein ensemble of parameterizations. Therefore, the proposed concurrent, condensed SVGD (ccSVGD) method can provide UQ on parameters, not just outputs. Furthermore, the parameter reduction speeds up the convergence of the Stein gradient descent as it reduces the combinatorial complexity by aligning and differentiating the sensitivity to parameters. These properties are demonstrated with an illustrative example and an application to a mechanical response representation problem in solid mechanics.

42 ENGINEERING↗