Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Stochastic inference”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Producing High-fidelity Synthetic Population Ensembles at Scale

Used within social simulations, synthetic population ensembles enable uncertainty quantification (UQ) methods for obtaining more robust model inference and prediction. A synthetic population ensemble is a series of plausible virtual reconstructions of an area’s population at the granularity of people and residences, generated stochastically to preserve privacy of the source population survey’s respondents. In this paper, we demonstrate the production of large synthetic population ensembles for the US via Oak Ridge National Laboratory’s UrbanPop framework to support modeling of high spatial resolution energy affordability metrics from nationwide social surveys in collaboration with the fusionACS project. Our initial task involves creating ensembles for 17 US metropolitan areas, each consisting of 41 population instances (a base realization and 40 replicates). To accomplish this task at scale, we configured an integrated system comprised of a research cloud, virtual containerization, GPU-enhanced functionality, and a dual API/CLI to interact with UrbanPop’s maturing Likeness Python ecosystem. We observe a reduction in theoretical execution time while maintaining high-fidelity approximations of residential totals by metropolitan area and the demographic characteristics of neighborhoods. We discuss expansion of our approach to produce synthetic population ensembles for the entire US, particularly plans to establish automated workflows for job orchestration to increase computational efficiency, as well as provide outlook for broadening applications of the ensembles.

Gaboardi, James [ORNL] (ORCID:0000000247766826)↗

The clustering of DESI-like luminous red galaxies using photometric redshifts

ABSTRACT We present measurements of the redshift-dependent clustering of a DESI-like luminous red galaxy (LRG) sample selected from the Legacy Survey imaging data set, and use the halo occupation distribution (HOD) framework to fit the clustering signal. The photometric LRG sample in this study contains 2.7 million objects over the redshift range of 0.4 < z < 0.9 over 5655 deg2. We have developed new photometric redshift (photo-z) estimates using the Legacy Survey DECam and WISE photometry, with σNMAD = 0.02 precision for LRGs. We compute the projected correlation function using new methods that maximize signal-to-noise ratio while incorporating redshift uncertainties. We present a novel algorithm for dividing irregular survey geometries into equal-area patches for jackknife resampling. For a five-parameter HOD model fit using the MultiDark halo catalogue, we find that there is little evolution in HOD parameters except at the highest redshifts. The inferred large-scale structure bias is largely consistent with constant clustering amplitude over time. In an appendix, we explore limitations of Markov chain Monte Carlo fitting using stochastic likelihood estimates resulting from applying HOD methods to N-body catalogues, and present a new technique for finding best-fitting parameters in this situation. Accompanying this paper, we have released the Photometric Redshifts for the Legacy Surveys catalogue of photo-z’s obtained by applying the methods used in this work to the full Legacy Survey Data Release 8 data set. This catalogue provides accurate photometric redshifts for objects with z < 21 over more than 16 000 deg2 of sky.

79 ASTRONOMY AND ASTROPHYSICS↗

Laser Wakefield Accelerator modelling with Variational Neural Networks

A machine learning model was created to predict the electron spectrum generated by a GeV-class laser wakefield accelerator. The model was constructed from variational convolutional neural networks, which mapped the results of secondary laser and plasma diagnostics to the generated electron spectrum. An ensemble of trained networks was used to predict the electron spectrum and to provide an estimation of the uncertainty of that prediction. It is anticipated that this approach will be useful for inferring the electron spectrum prior to undergoing any process that can alter or destroy the beam. In addition, the model provides insight into the scaling of electron beam properties due to stochastic fluctuations in the laser energy and plasma electron density.

43 PARTICLE ACCELERATORS↗

Patterns, drivers, and a predictive model of dam removal cost in the United States

Given the burgeoning dam removal movement and the large number of dams approaching obsolescence in the United States, cost estimating data and tools are needed for dam removal prioritization, planning, and execution. We used the list of removed dams compiled by American Rivers to search for publicly available reported costs for dam removal projects. Total cost information could include component costs related to project planning, dam deconstruction, monitoring, and several categories of mitigation activities. We compiled reported costs from 455 unique sources for 668 dams removed in the United States from 1965 to 2020. The dam removals occurred within 571 unique projects involving 1–18 dams. When adjusted for inflation into 2020 USD, cost of these projects totaled $\$1.522$ billion, with per-dam costs ranging from $\$1$ thousand (k) to $\$268.8$ million (M). The median cost for dam removals was $\$157$k, $\$823$k, and $\$6.2$M for dams that were< 5 m, between 5–10 m, and > 10 m in height, respectively. Geographic differences in total costs showed that northern states in general, and the Pacific Northwest in particular, spent the most on dam removal. The Midwest and the Northeast spent proportionally more on removal of dams less than 5 m in height, whereas the Northwest and Southwest spent the most on larger dam removals > 10 m tall. We used stochastic gradient boosting with quantile regression to model dam removal cost against potential predictor variables including dam characteristics (dam height and material), hydrography (average annual discharge and drainage area), project complexity (inferred from construction and sediment management, mitigation, and post-removal cost drivers), and geographic region. Dam height, annual average discharge at the dam site, and project complexity were the predominant drivers of removal cost. The final model had an R 2 of 57% and when applied to a test dataset model predictions had a root mean squared error of $\$5.09$M and a mean absolute error of $\$1.45$M, indicating its potential utility to predict estimated costs of dam removal. We developed a R shiny application for estimating dam removal costs using customized model inputs for exploratory analyses and potential dam removal planning.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

UQ Toolkit v 2.0

The Uncertainty Quantification (UQ) Toolkit is a software library for the characterizaton and propagation of uncertainties in computational models. For the characterization of uncertainties, Bayesian inference tools are provided to infer uncertain model parameters, as well as Bayesian compressive sensing methods for discovering sparse representations of high-dimensional input-output response surfaces, and also Karhunen-Loève expansions for representing stochastic processes. Uncertain parameters are treated as random variables and represented with Polynomial Chaos expansions (PCEs). The library implements several spectral basis function types (e.g. Hermite basis functions in terms of Gaussian random variables or Legendre basis functions in terms of uniform random variables) that can be used to represent random variables with PCEs. For propagation of uncertainty, tools are provided to propagate PCEs that describe the input uncertainty through the computational model using either intrusive methods (Galerkin projection of equations onto basis functions) or non-intrusive methods (perform deterministic operation at sampled values of the random values and project the obtained results onto basis functions).

Safta, Cosmin↗

The Mock LISA Data Challenge Round 3: New and Improved Sources

The Mock LISA Data Challenges are a program to demonstrate and encourage the development of data-analysis capabilities for LISA. Each round of challenges consists of several data sets containing simulated instrument noise and gravitational waves from sources of undisclosed parameters. Participants are asked to analyze the data sets and report the maximum information they can infer about the source parameters. The challenges are being released in rounds of increasing complexity and realism. Challenge 3. currently in progress, brings new source classes, now including cosmic-string cusps and primordial stochastic backgrounds, and more realistic signal models for supermassive black-hole inspirals and galactic double white dwarf binaries.

Baker, John↗

Stochastic Reconstruction of Thermal Protection Material Properties from Arc-Jet Experiments

Material response models are used to assess reliability using variances in the bond-line temperature predictions based on uncertainties in trajectory, aerothermal environment, and material properties. A key deficiency in the current approach is that input uncertainties are too often subjective, empirical, or ad-hoc, and are not rigorously linked to the arc-jet test data used to develop the TPS material model. While materials such as PICA are well understood, future missions may require more novel materials such as HEEET where unknown uncertainties have real consequences on the ability to assess reliability. A quantifiable estimate of reliability requires an iterative methodology where the parameters driving the variance in bond-line temperature (for example) are systematically identified. A test campaign to collect data or develop new models can then be identified to reduce those input uncertainties. A Bayesian inference loop defines these connections mathematically, i.e., prior knowledge about uncertainty is updated based on observation. While these concepts are well known (and often applied intuitively in a non-rigorous approach), only recent advances in reduced-order modelling have made them computationally viable methods for engineering. By replacing deterministic inverse methods with stochastic approaches, the hope is new materials proposed for future missions can more rapidly be developed with a greater understanding of the TPS material reliability. Two additional steps for the analysis of arc jet test data are discussed. The first is ability to construct a reduced-order model using material response simulations (Icarus/US3D) of the arc-jet test articles, and the second is the inclusion of this surrogate model in the Bayesian inversion process. Both capabilities will be demonstrated using prior PICA arc-jet test data. The quality of a surrogate model will be investigated and the variances on the calibrated material properties will be compared to our current understanding of the PICA material model.

Material response↗

Combined selection of the dynamic model and modeling error in nonlinear aeroelastic systems using Bayesian Inference

Here, we report a Bayesian framework for concurrent selection of physics-based models and (modeling) error models. We investigate the use of colored noise to capture the mismatch between the predictions of calibrated models and observational data that cannot be explained by measurement error alone within the context of Bayesian estimation for stochastic ordinary differential equations. Proposed models are characterized by the average data-fit, a measure of how well a model fits the measurements, and the model complexity measured using the Kullback–Leibler divergence. The use of a more complex error models increases the average data-fit but also increases the complexity of the combined model, possibly over-fitting the data. Bayesian model selection is used to find the optimal physical model as well as the optimal error model. The optimal model is defined using the evidence, where the average data-fit is balanced by the complexity of the model. The effect of colored noise process is illustrated using a nonlinear aeroelastic oscillator representing a rigid NACA0012 airfoil undergoing limit cycle oscillations due to complex fluid–structure interactions. Several quasi-steady and unsteady aerodynamic models are proposed with colored noise or white noise for the model error. The use of colored noise improves the predictive capabilities of simpler models.

42 ENGINEERING↗

Reconstruction of effective potential from statistical analysis of dynamic trajectories

The broad incorporation of microscopic methods is yielding a wealth of information on the atomic and mesoscale dynamics of individual atoms, molecules, and particles on surfaces and in open volumes. Analysis of such data necessitates statistical frameworks to convert observed dynamic behaviors to effective properties of materials. Here, we develop a method for the stochastic reconstruction of effective local potentials solely from observed structural data collected from molecular dynamics simulations (i.e., data analogous to those obtained via atomically resolved microscopies). Using the silicon vacancy defect in graphene as a model, we apply the statistical framework presented herein to reconstruct the free energy landscape from the calculated atomic displacements. Evidence of consistency between the reconstructed local potential and the trajectory data from which it was produced is presented, along with a quantitative assessment of the uncertainty in the inferred parameters.

74 ATOMIC AND MOLECULAR PHYSICS↗

HyDE Framework for Stochastic and Hybrid Model-Based Diagnosis

Hybrid Diagnosis Engine (HyDE) is a general framework for stochastic and hybrid model-based diagnosis that offers flexibility to the diagnosis application designer. The HyDE architecture supports the use of multiple modeling paradigms at the component and system level. Several alternative algorithms are available for the various steps in diagnostic reasoning. This approach is extensible, with support for the addition of new modeling paradigms as well as diagnostic reasoning algorithms for existing or new modeling paradigms. HyDE is a general framework for stochastic hybrid model-based diagnosis of discrete faults; that is, spontaneous changes in operating modes of components. HyDE combines ideas from consistency-based and stochastic approaches to model- based diagnosis using discrete and continuous models to create a flexible and extensible architecture for stochastic and hybrid diagnosis. HyDE supports the use of multiple paradigms and is extensible to support new paradigms. HyDE generates candidate diagnoses and checks them for consistency with the observations. It uses hybrid models built by the users and sensor data from the system to deduce the state of the system over time, including changes in state indicative of faults. At each time step when observations are available, HyDE checks each existing candidate for continued consistency with the new observations. If the candidate is consistent, it continues to remain in the candidate set. If it is not consistent, then the information about the inconsistency is used to generate successor candidates while discarding the candidate that was inconsistent. The models used by HyDE are similar to simulation models. They describe the expected behavior of the system under nominal and fault conditions. The model can be constructed in modular and hierarchical fashion by building component/subsystem models (which may themselves contain component/ subsystem models) and linking them through shared variables/parameters. The component model is expressed as operating modes of the component and conditions for transitions between these various modes. Faults are modeled as transitions whose conditions for transitions are unknown (and have to be inferred through the reasoning process). Finally, the behavior of the components is expressed as a set of variables/ parameters and relations governing the interaction between the variables. The hybrid nature of the systems being modeled is captured by a combination of the above transitional model and behavioral model. Stochasticity is captured as probabilities associated with transitions (indicating the likelihood of that transition being taken), as well as noise on the sensed variables.

Narasimhan, Sriram↗

Calibration verification for stochastic agent-based disease spread models

Accurate disease spread modeling is crucial for identifying the severity of outbreaks and planning effective mitigation efforts. To be reliable when applied to new outbreaks, model calibration techniques must be robust. However, current methods frequently forgo calibration verification (a stand-alone process evaluating the calibration procedure) and instead use overall model validation (a process comparing calibrated model results to data) to check calibration processes, which may conceal errors in calibration. In this work, we develop a stochastic agent-based disease spread model to act as a testing environment as we test two calibration methods using simulation-based calibration, which is a synthetic data calibration verification method. The first calibration method is a Bayesian inference approach using an empirically-constructed likelihood and Markov chain Monte Carlo (MCMC) sampling, while the second method is a likelihood-free approach using approximate Bayesian computation (ABC). Simulation-based calibration suggests that there are challenges with the empirical likelihood calculation used in the first calibration method in this context. These issues are alleviated in the ABC approach. Despite these challenges, we note that the first calibration method performs well in a synthetic data model validation test similar to those common in disease spread modeling literature. We conclude that stand-alone calibration verification using synthetic data may benefit epidemiological researchers in identifying model calibration challenges that may be difficult to identify with other commonly used model validation techniques.

60 APPLIED LIFE SCIENCES↗

Joint Modeling of Quasar Variability and Accretion Disk Reprocessing Using Latent Stochastic Differential Equations

Quasars are bright active galactic nuclei powered by the accretion of matter around supermassive black holes at the center of galaxies. Their stochastic brightness variability depends on the physical properties of the accretion disk and black hole. The upcoming Rubin Observatory Legacy Survey of Space and Time (LSST) is expected to observe tens of millions of quasars, so there is a need for efficient techniques like machine learning that can handle the large volume of data. Quasar variability is believed to be driven by an X-ray corona, which is reprocessed by the accretion disk and emitted as UV/optical variability. We are the first to introduce an auto-differentiable simulation of the accretion disk and reprocessing. We use the simulation as a direct component of our neural network to jointly model the driving variability and reprocessing, trained with supervised learning on simulated LSST-like 10 yr quasar light curves. We encode the light curves using a transformer encoder, and the driving variability is reconstructed using latent stochastic differential equations, a physically motivated generative deep learning method that can model continuous-time stochastic dynamics. By embedding the physical processes of the driving signal and reprocessing into our network, we achieve a model that is more robust and interpretable. We demonstrate that our model outperforms a Gaussian process regression baseline and can infer accretion disk parameters and time delays between wave bands, even for out-of-distribution driving signals. Our approach provides a powerful framework that can be adapted to solve other inverse problems in multivariate time series.

Fagin, Joshua [City Univ. of New York (CUNY), NY (↗

Portfolios in Stochastic Local Search: Efficiently Computing Most Probable Explanations in Bayesian Networks

Portfolio methods support the combination of different algorithms and heuristics, including stochastic local search (SLS) heuristics, and have been identified as a promising approach to solve computationally hard problems. While successful in experiments, theoretical foundations and analytical results for portfolio-based SLS heuristics are less developed. This article aims to improve the understanding of the role of portfolios of heuristics in SLS. We emphasize the problem of computing most probable explanations (MPEs) in Bayesian networks (BNs). Algorithmically, we discuss a portfolio-based SLS algorithm for MPE computation, Stochastic Greedy Search (SGS). SGS supports the integration of different initialization operators (or initialization heuristics) and different search operators (greedy and noisy heuristics), thereby enabling new analytical and experimental results. Analytically, we introduce a novel Markov chain model tailored to portfolio-based SLS algorithms including SGS, thereby enabling us to analytically form expected hitting time results that explain empirical run time results. For a specific BN, we show the benefit of using a homogenous initialization portfolio. To further illustrate the portfolio approach, we consider novel additive search heuristics for handling determinism in the form of zero entries in conditional probability tables in BNs. Our additive approach adds rather than multiplies probabilities when computing the utility of an explanation. We motivate the additive measure by studying the dramatic impact of zero entries in conditional probability tables on the number of zero-probability explanations, which again complicates the search process. We consider the relationship between MAXSAT and MPE, and show that additive utility (or gain) is a generalization, to the probabilistic setting, of MAXSAT utility (or gain) used in the celebrated GSAT and WalkSAT algorithms and their descendants. Utilizing our Markov chain framework, we show that expected hitting time is a rational function - i.e. a ratio of two polynomials - of the probability of applying an additive search operator. Experimentally, we report on synthetically generated BNs as well as BNs from applications, and compare SGSs performance to that of Hugin, which performs BN inference by compilation to and propagation in clique trees. On synthetic networks, SGS speeds up computation by approximately two orders of magnitude compared to Hugin. In application networks, our approach is highly competitive in Bayesian networks with a high degree of determinism. In addition to showing that stochastic local search can be competitive with clique tree clustering, our empirical results provide an improved understanding of the circumstances under which portfolio-based SLS outperforms clique tree clustering and vice versa.

Mengshoel, Ole J.↗

Correlation-aware binning for small-angle neutron scattering via Gaussian-process inference

Binning in small-angle neutron scattering (SANS) is typically performed empirically, with fixed parameters chosen for convenience rather than statistical optimality. Such practices often fail to balance statistical precision and spatial resolution, leading to inconsistencies across instruments and datasets. Here we establish a correlation-aware framework that determines the optimal bin width from first principles by extending the classical Freedman–Diaconis (FD) rule to account for inter-bin correlations with a Gaussian process. In this formulation, the scattering intensity is treated as a smooth stochastic field whose statistical coherence is described by a covariance matrix. Analytical expressions of errors derived from this model yield closed-form criteria that separate the total deviation into contributions from counting noise, aliasing distortion and curvature-dependent correlation effects. Expressed in reduced variables, the resulting dimensionless error surface reveals a continuous transition from the uncorrelated FD regime to the correlation-dominated limit, providing a unified description of noise suppression and resolution control. Because the formulation depends only on the profile characteristics of scattering intensity I(Q), specifically its average intensity and first- and second-order derivatives, it applies generally to any SANS measurement regardless of sample, instrument or geometry. Experimental validation using small- and ultra-small-angle neutron scattering data confirms the predicted scaling behavior, demonstrating that correlation-aware inference systematically reduces mean-squared error and enables information-efficient reproducible data reduction across materials and instruments.

Tung, Chi-Huan [ORNL] (ORCID:0000000221972074)↗

Spatiotemporal 4D Whole-cell Modeling of a Minimal Autotroph Reveals Central Carbon Metabolism Regulated Locally by Protein Megacomplexes via Post-translational Modifications under Light Disturbance

Photosynthetic microorganisms rely on multiple pathways in central carbon metabolism to adapt to fluctuating light and energy availability across diel cycles. Mechanistic insight into the regulatory dynamics of this adaptation requires integrating processes spanning disparate timescales, from rapid redox-dependent post-translational modifications (PTMs) to slower changes in protein expression and metabolic pathway usage. To address this complexity beyond genome-based inference and traditional modeling, we develop a whole-cell four-dimensional (3D + time) model of the marine cyanobacterium Prochlorococcus marinus MED4 that explicitly represents the spatial organization of enzymatic and molecular processes in central carbon metabolism under light perturbation. We employ a perturbation-based research design to experimentally generate time-series, multi-omics measurements that provide molecular descriptors and cryo-ET derived 3D segmented volumes as constraints for this dynamic 4D framework. The integration of experiments and modeling across defined light regimes enables quantitative validation of system-level responses and forecasting under distinct light disturbances. We test the hypothesis that light-dependent redox PTMs regulating the structural assembly of a protein megacomplex, the “dark complex,” modulate metabolic flux at a conserved regulatory node of the Calvin–Benson cycle (CBC) in cyanobacteria. Our model shows that subcellular spatial organization buffers rapid light-induced changes in thylakoid reaction rates, which are followed by redox-PTM-mediated sequestration or release of CBC enzymes in the dark complex, ultimately impacting carbon fixation dynamics within carboxysomes. Comparison with an equivalently parameterized well-mixed stochastic model demonstrates that post-translational regulation not only buffers transcriptional noise and diffusion-driven fluctuations but also stabilizes phenotypic outcomes, underscoring the importance of spatial heterogeneity in phenotypic robustness. This ability to probe adaptive, spatiotemporally resolved mechanisms in photosynthetic machinery and central carbon metabolism addresses a critical gap in genotype-to-phenotype inference and expands modeling and design capabilities for understudied or genetically intractable autotrophs such as P. marinus MED4.

Johnson, Connah G.↗

Precipitating electron energy flux and auroral zone conductances - An empirical model

Data from the low energy electron (LEE) experiments on the Atmosphere Explorer C and D satellites have been used to determine the average global distribution of the energy flux of precipitating auroral electrons and their average energy for different levels of geomagnetic activity. Measurements from the Atmosphere Explorer unified abstract file (15-s resolution) have been binned according to invariant latitude (in the range 50-90 deg), magnetic local time, and geomagnetic activity as measured by the Kp and auroral electrojet (AE) indices, separately. Bin-averaged values of precipitating energy flux and average energy have been calculated, and a smoothing algorithm used to reduce stochastic variations in the raw data. The results indicate that, for the parameters studied, the AE inces does a superior job of ordering the data with regard to geomagnetic activity. The global distribution of the auroral enhancement porition of the Pedersen and Hall conductances were inferred from the data by means of an empirical fit to detailed energy deposition calculations.

Spiro, R. W.↗

Low-frequency mobility response functions for the central plasma sheet with application to tearing modes

Consideration is given to the effect of constant cross-tail magnetic field By on the collisionless conductivity produced by chaotic scattering and stochastic diffusion of particles in the current sheet for a parabolic geometry. It is shown that the correlation time scales as (By/Bz)-squared, and from this strong By scaling a strong tendency toward stabilization of the linear tearing modes with increasing values of By is inferred. This effect of increased dawn-dusk mobility is particularly dramatic when electrons are introduced in the calculation, and is in agreement with the results of kinetic particle simulations. The collisionless conductivity is expressed in terms of the ensemble-averaged power spectrum of the single particle trajectories, which makes it possible to calculate directly the linear conductivity instead of deriving it from the calculation of the irreversible heating rates.

Hernandez, J.↗