Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Random fields”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Text Mining for Process–Structure–Properties Relationships in Metals

With the advent of large language models (LLMs), the vast unstructured text within millions of academic papers is increasingly accessible for materials discovery—although significant challenges remain. While LLMs offer promising few- and zero-shot learning capabilities, particularly valuable in the materials domain where expert annotations are scarce, general-purpose LLMs often fail to address key materials-specific queries without further adaptation. To bridge this gap, fine-tuning LLMs on human-labeled data is essential for effective structured knowledge extraction (Liu in The Importance of Human-Labeled Data in the Era of LLMs, 2023). Here, in this study, we introduce a novel annotation schema designed to extract generic process–structure–properties relationships from scientific literature. We demonstrate the utility of this approach using a dataset of 128 abstracts, with annotations drawn from two distinct domains: high-temperature materials (Domain I) and uncertainty quantification in simulating materials microstructure (Domain II). Initially, we developed a conditional random field (CRF) model based on MatBERT—a domain-specific BERT variant—and evaluated its performance on Domain I. Subsequently, we compared this model with a fine-tuned LLM (GPT-4o from OpenAI) under identical conditions. Our results indicate that fine-tuning LLMs can significantly improve entity extraction performance over the BERT-CRF baseline on Domain I. However, when additional examples from Domain II were incorporated, the performance of the BERT-CRF model became comparable to that of the GPT-4o model. These findings underscore the potential of our schema for structured knowledge extraction and highlight the complementary strengths of both modeling approaches.

Materials science↗

Estimation of hydraulic conductivity in a watershed using sparse multi-source data via Gaussian process regression and Bayesian experimental design

Enhanced water management systems depend on accurate estimation of subsurface hydraulic properties. However, geologic formations can vary significantly, so information from a single source (e.g., widely spaced boreholes) is insufficient in characterizing subsurface aquifer properties. Therefore, multiple sources of information are needed to complement the hydrogeology understanding of a region. Here, this study presents a numerical framework in which information from different measurement sources is combined to characterize the 3D random field in a multi-fidelity prediction model. Coupled with the model, a Bayesian experimental design was used to determine the best future sampling locations. The Upper Sangamon watershed in east-central Illinois was selected as the case study site, where the multi-fidelity Gaussian process model was used to estimate the hydraulic conductivity in the region of interest. Multi-source observation data were obtained from electrical resistivity and borehole pumping tests. The accuracy of the model prediction is dependent on the locations and the distribution of both high- and low-fidelity data. Furthermore, the multi-fidelity model was compared with the single-fidelity model. The uncertainties and confidence in the measurements and parameter estimates were quantified and used to design future cycles of data collection to further improve the confidence intervals.

54 ENVIRONMENTAL SCIENCES↗

Hierarchical reconstruction of 3D well-connected porous media from 2D exemplars using statistics-informed neural network

The relationships between porous microstructures and transport properties are of fundamental importance in various scientific and engineering applications. Due to the intricacy, stochasticity and heterogeneity of porous media, reliable characterization and modeling of transport properties often require a complete dataset of internal microstructure samples. However, it is often an unbearable cost to acquire sufficient 3D digital microstructures by purely using microscopic imaging systems. Herein this paper presents a machine learning-based technique to hierarchically reconstruct 3D well-connected porous microstructures from one isotropic or several anisotropic low-cost 2D exemplar(s). To compactly characterize the large-scale microstructural features, a Gaussian image pyramid is built for each 2D exemplar. Local morphology patterns are collected from the Gaussian image pyramids, and then they serve as the training data to embed the 2D morphological statistics into feed-forward neural networks at multiple length levels. By using a specially-developed morphology integration scheme, the 3D morphological statistics at different levels can be inferred from the statistics-informed neural networks. Gibbs sampling is adopted to hierarchically reconstruct 3D microstructures by using multi-level 3D morphological statistics, where the large-scale, regional and local morphological patterns are statistically generated and successively added to the same 3D random field. The proposed method is tested on a series of porous media with distinct morphologies, and the statistical equivalence between the reconstructed and the real microstructures is systematically evaluated by comparing morphological descriptors and transport properties. The results demonstrate that the proposed 2D-to-3D microstructure reconstruction method is a universal and efficient approach to generating morphologically and physically realistic samples of porous media.

42 ENGINEERING↗

A better understanding of the mechanics of borehole breakout utilizing a finite strain gradient-enhanced micropolar continuum model

Borehole breakout denotes the failure in rock mass subjected to drilling, caused by stress concentrations exceeding the material strength. Depending on the material properties, the preexisting in situ stress state, and the borehole dimensions, different types of borehole breakout, such as spiral-shaped breakout or v-shaped breakout, are distinguished in the literature. In the present work, we address the influence of the material properties on the predicted borehole breakout mode in a comprehensive finite element study. To this end, herein we employ a gradient-enhanced micropolar damage-plasticity model based on the Mohr–Coulomb strength criterion, formulated in the finite strain regime, which is calibrated based on experimental results from plane strain compression tests on Red Wildmoor sandstone. In the numerical study, the influence of the in situ stress state, the material friction angle, plastic dilation, post peak residual strength, and the inherent material length scale parameters are investigated. Thereby, we demonstrate that depending on the material parameters, substantially different failure modes, characterized by strongly localized shear bands or diffuse failure zones, are predicted. It is shown that in particular the brittleness of the material in the post peak regime has a major influence on the predicted breakout type. Moreover, a statistical validation of the results is obtained by considering different random field distributions of the initial material strength.

58 GEOSCIENCES↗

Emulation and detection of physical faults and cyber-attacks on building energy systems through real-time hardware-in-the-loop experiments

The increasing use of remote or mobile access, integrated wearable technologies, data exchange, and cloud-based data analytics in modern smart buildings is steering the building industry towards open communication technologies. The increased connectivity and accessibility could lead to more cyber-attacks in smart buildings. On the other hand, physical faults (e.g., HVAC -heating, ventilation, and air-conditioning faults) may have similar adverse impacts as those from the cyber-attacks on building energy systems, such as occupant discomfort, energy wastage, and equipment downtime. However, current physical behavior-based anomaly detection methods fail to differentiate between cyber-attacks and physical faults in building energy systems. Moreover, the challenge in collecting real-world threat data with ground truth has led researchers to rely on numerical models with user-defined assumptions, which may not accurately reflect real-world conditions due to the lack of in-situ experimental datasets. To address these challenges and gaps, this paper presents a flexible hardware-in-the-loop (HIL) testbed for generating cyber-attack and physical fault datasets and demonstrating threat detection algorithms in a real building automation system (BAS) environment. This testbed combines hardware (i.e., real BAS with local HVAC controllers and a physical network) with software (i.e., high-fidelity models to represent behaviors of building envelope and HVAC energy systems), enabling emulations of realistic threats. Five HIL experiments, including one baseline without any threats, two with physical faults, and two with cyber-attacks, were conducted to generate datasets containing detailed network traffic and system states. A joint classification framework, incorporating a network analyzer and a physical HVAC fault detector, was proposed to automatically detect cyber-physical abnormalities on BAS at both the network and the physical HVAC levels. The network analyzer comprises a conditional random fields (CRF) based command validator and a statistics-based detection strategy. The fault detector employs a weather and schedule-based pattern matching and feature-based principal component analysis (WPM-FPCA) method. Evaluation of the classification using four metrics from the multi-class confusion matrix revealed an average accuracy of 90.2%, recall of 89.7%, precision of 88.5% and F1-score of 89.2%. Finally, these results demonstrate that the proposed joint classification framework can effectively differentiate between specific types of cyber-attacks (e.g., device reinitialization attack, network Denial-of-Service attack) and physical faults (e.g., air handling unit operational fault, cooling coil valve stuck) in real time for improved building energy management.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A scalable planning framework of energy storage systems under frequency dynamics constraints

As the penetration of renewables increases in power systems, the declining system inertia can cause frequency stability issues. Battery energy storage systems (BESSs) respond fast and therefore can relieve the low inertia difficulty but need to be appropriately sized considering the associated cost. This paper presents a novel stochastic optimization model for economically planning BESS capacity while considering the spatial–temporal correlation of wind generation and generator outages under frequency stability constraints, which include the rate-of-change of frequency (RoCoF), frequency nadir (FN), and quasi-steady-state (QSS) frequency. A set of new FN constraints that can be easily linearized is developed. To account for renewable uncertainties, a realistic uncertainty modeling approach, Random Field, is adopted to generate wind generation scenarios by considering both spatial and temporal evolutions of wind speed profiles. The ESS sizing is formulated as a mixed-integer linear programming problem and solved by using a scalable decomposition-and-coordination approach, Surrogate Absolute Value Lagrangian Relaxation (SAVLR). To further improve the scalability and reduce computational burdens, a rolling-horizon-based update is developed and incorporated into SAVLR for providing a practical solution to the long-term planning of very large-scale power systems. Finally, a modified IEEE 118-bus system and the Polish system are used to validate the effectiveness and scalability of the model and solution methodology.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Exploring drought-responsive crucial genes in Sorghum

Drought severely affects global food production. Sorghum is a typical drought-resistant model crop. Based on RNA-seq data for Sorghum with multiple time points and the gray correlation coefficient, this paper firstly selects candidate genes via mean variance test and constructs weighted gene differential co-expression networks (WGDCNs); then, based on guilt-by-rewiring principle, the WGDCNs and the hidden Markov random field model, drought-responsive crucial genes are identified for five developmental stages respectively. Enrichment and sequence alignment analysis reveal that the screened genes may play critical functional roles in drought responsiveness. A multilayer differential co-expression network for the screened genes reveals that Sorghum is very sensitive to pre-flowering drought. Furthermore, a crucial gene regulatory module is established, which regulates drought responsiveness via plant hormone signal transduction, MAPK cascades, and transcriptional regulations. The proposed method can well excavate crucial genes through RNA-seq data, which have implications in breeding of new varieties with improved drought tolerance.

60 APPLIED LIFE SCIENCES↗

A survey of unsupervised learning methods for high-dimensional uncertainty quantification in black-box-type problems

Constructing surrogate models for uncertainty quantification (UQ) on complex partial differential equations (PDEs) having inherently high-dimensional O(10 n ), n ≥ 2, stochastic inputs (e.g., forcing terms, boundary conditions, initial conditions) poses tremendous challenges. The “curse of dimensionality” can be addressed with suitable unsupervised learning techniques used as a pre-processing tool to encode inputs onto lower-dimensional subspaces while retaining its structural information and meaningful properties. In this work, we review and investigate thirteen dimension reduction methods including linear and nonlinear, spectral, blind source separation, convex and non-convex methods and utilize the resulting embeddings to construct a mapping to quantities of interest via polynomial chaos expansions (PCE). Here, we refer to the general proposed approach as manifold PCE (m-PCE), where manifold corresponds to the latent space resulting from any of the studied dimension reduction methods. To investigate the capabilities and limitations of these methods we conduct numerical tests for three physics-based systems (treated as black-boxes) having high-dimensional stochastic inputs of varying complexity modeled as both Gaussian and non-Gaussian random fields to investigate the effect of the intrinsic dimensionality of input data. We demonstrate both the advantages and limitations of the unsupervised learning methods and we conclude that a suitable m-PCE model provides a cost-effective approach compared to alternative algorithms proposed in the literature, including recently proposed expensive deep neural network-based surrogates and can be readily applied for high-dimensional UQ in stochastic PDEs.

42 ENGINEERING↗

Mesh objective stochastic simulations of quasibrittle fracture

Continuum finite element (FE) modeling of damage and failure of quasibrittle structures suffers from the spurious mesh sensitivity due to strain localization. Here this issue has been addressed for deterministic analysis through the development of localization limiters. Here this study proposes a mechanism-based model to mitigate the mesh sensitivity in stochastic FE simulations of quasibrittle fracture. The interest is placed on the analysis of large-size structures, where the mesh size is conveniently chosen to be larger than the width of the fracture process zone as well as the correlation length of the random fields of constitutive properties. The present model is formulated within the framework of continuum damage mechanics. Two localization parameters are introduced to describe the evolution of the damage pattern of each finite element. These parameters are used to guide the energy regularization of the constitutive law, as well as to formulate the mesh-dependent probability distributions of constitutive properties. Depending on the prevailing damage pattern, different energy regularization schemes and mesh dependence of the probability distribution functions are used in the constitutive law. The model is applied to simulate the stochastic failure behavior of quasibrittle structures of different geometries featuring different failure processes including damage initiation, localization, and propagation. It is shown that using fixed probability distribution functions of constitutive properties could lead to strong mesh dependence of the prediction of the mean and variance of the structural load capacity. The probability distribution functions of constitutive properties must be linked to the damage pattern, which may evolve during the failure process. Such a mechanism-based modeling of the probability distributions of constitutive properties is essential for mitigating the spurious mesh sensitivity in stochastic FE analysis of quasibrittle fracture.

42 ENGINEERING↗

A Bayesian model for multivariate discrete data using spatial and expert information with application to inferring building attributes

When modeling sparsely observed multivariate data, strong prior information elicited from experts can be used to bolster predictive accuracy and counteract sampling bias. Similarly, modeling autocorrelation in space can help make use of co-occurrence patterns present in many types of spatial data. To make use of both expert prior information and spatial structure, we propose a novel graphical model for a spatial Bayesian network developed specifically to address challenges in inferring the attributes of buildings from geographically sparse observational data. This model is implemented as the sum of a spatial multivariate Gaussian random field and a tabular conditional probability function in real-valued space prior to projection onto the probability simplex. This modeling form is especially suitable for the usage of prior information in the form of sets of atomic rules obtained from experts. To perform inference with missing data, we implement a Markov chain Monte Carlo scheme composed of alternating steps of Gibbs sampling of missing entries and Hamiltonian Monte Carlo for model parameters. A case study in building attribution is presented to highlight the advantages and limitations of this approach.

97 MATHEMATICS AND COMPUTING↗

Critical nematic correlations throughout the superconducting doping range in Bi 2–z Pb z Sr 2–y La y CuO 6+x

Charge modulations have been widely observed in cuprates, suggesting their centrality for understanding the high-T c superconductivity in these materials. However, the dimensionality of these modulations remains controversial, including whether their wavevector is unidirectional or bidirectional, and also whether they extend seamlessly from the surface of the material into the bulk. Material disorder presents severe challenges to understanding the charge modulations through bulk scattering techniques. We use a local technique, scanning tunneling microscopy, to image the static charge modulations on Bi 2–z Pb z Sr 2–y La y CuO 6+x . The ratio of the phase correlation length ξ CDW to the orientation correlation length ξ orient points to unidirectional charge modulations. By computing new critical exponents at free surfaces including that of the pair connectivity correlation function, we show that these locally 1D charge modulations are actually a bulk effect resulting from classical 3D criticality of the random field Ising model throughout the entire superconducting doping range.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Identification of mobile genetic elements with geNomad

Identifying and characterizing mobile genetic elements in sequencing data is essential for understanding their diversity, ecology, biotechnological applications and impact on public health. Here we introduce geNomad, a classification and annotation framework that combines information from gene content and a deep neural network to identify sequences of plasmids and viruses. geNomad uses a dataset of more than 200,000 marker protein profiles to provide functional gene annotation and taxonomic assignment of viral genomes. Using a conditional random field model, geNomad also detects proviruses integrated into host genomes with high precision. In benchmarks, geNomad achieved high classification performance for diverse plasmids and viruses (Matthews correlation coefficient of 77.8% and 95.3%, respectively), substantially outperforming other tools. Leveraging geNomad’s speed and scalability, we processed over 2.7 trillion base pairs of sequencing data, leading to the discovery of millions of viruses and plasmids that are available through the IMG/VR and IMG/PR databases. geNomad is available at https://portal.nersc.gov/genomad.

59 BASIC BIOLOGICAL SCIENCES↗

Active and passive defects in tetragonal tungsten bronze relaxor ferroelectrics

Tetragonal tungsten bronze (TTB) based oxides constitute a large family of dielectric materials which are known to exhibit complex distortions producing incommensurately modulated superstructures as well as significant local deviations from their average symmetry. The local deviations produce diffuse scattering in diffraction experiments. The structure as well as the charge dynamics of these materials are anticipated to be sensitive to defects, such as cation or oxygen vacancies. In this work, in an effort to understand how the structural and charge dynamical properties respond to these two types of vacancy defects, we have performed measurements of dielectric susceptibilities and single crystal diffraction experiments of two types of TTB materials with both ‘filled’ (Ba 2 NdFeNb 4 O 15 and Ba 2 PrFeNb 4 O 15 ) and ‘unfilled’ (Sr 0.5 Ba 0.5 Nb 2 O 6 ) cation sublattices. We also perform these measurements before and after oxygen annealing, which alters the oxygen vacancy concentrations. Surprisingly, we find that many of the diffuse scattering features that are present in the unfilled structure are also present in the filled structure, suggesting that the random fields and disorder that are characteristic of the unfilled structure are not responsible for many of the local structural features that are reflected in the diffuse scattering. Furthermore, oxygen annealing clearly affected both color and dielectric properties, consistent with a diminishment of the oxygen vacancy concentration, but had little effect on observed diffuse patterns.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Mock data sets for the Eboss and DESI Lyman-α forest surveys

We present a publicly-available code to generate sets of mock Lyman-α (Lyα) forest data that have realistic large-scale correlations including those due to the Baryonic Acoustic Oscillations (BAO). The primary purpose of these mocks is to test the analysis procedures of the Extended Baryon Oscillation Survey (eBOSS) and the Dark Energy Spectroscopy Instrument (DESI) surveys. The transmitted flux fraction, F(λ), of background quasars due to Lyα absorption in the intergalactic medium (IGM) is simulated using the Fluctuating Gunn-Petterson Approximation (FGPA) applied to Gaussian random fields produced through the use of fast Fourier transforms (FFT). The output includes the IGM-Lyα transmitted flux fraction along quasar lines of sight and a catalog of high-column-density systems appropriately placed at high-density regions of the IGM. This output serves as input to additional code that superimposes the IGM tranmission on realistic quasar spectra, adds absorption by high-column-density systems and metals, and simulates instrumental transmission and noise. Redshift space distortions (RSD) of the flux correlations are implemented by including the large-scale velocity-gradient field in the FGPA resulting in a correlation function of F(λ) that can be accurately predicted. One hundred realizations have been produced over the 14,000 deg 2 DESI survey footprint with 100 quasars per deg 2 . The analysis of these realizations shows that the correlations of F(λ) follows the prediction within the accuracy of eBOSS survey. Here, the most time-consuming part of the mock production occurs before application of the FGPA, and the existing pre-FGPA forests can be used to easily produce new mock sets with modified redshift-dependent bias parameters or observational conditions.

79 ASTRONOMY AND ASTROPHYSICS↗

Validation of the DESI 2024 Lyα forest BAO analysis using synthetic datasets

The first year of data from the Dark Energy Spectroscopic Instrument (DESI) contains the largest set of Lyman-α (Lyα) forest spectra ever observed. This data, collected in the DESI Data Release 1 (DR1) sample, has been used to measure the Baryon Acoustic Oscillation (BAO) feature at redshift z = 2.33. In this work, we use a set of 150 synthetic realizations of DESI DR1 to validate the DESI 2024 Lyα forest BAO measurement presented in [1]. The synthetic data sets are based on Gaussian random fields using the log-normal approximation. We produce realistic synthetic DESI spectra that include all major contaminants affecting the Lyα forest. The synthetic data sets span a redshift range 1.8 < z < 3.8, and are analyzed using the same framework and pipeline used for the DESI 2024 Lyα forest BAO measurement. To measure BAO, we use both the Lyα auto-correlation and its cross-correlation with quasar positions. We use the mean of correlation functions from the set of DESI DR1 realizations to show that our model is able to recover unbiased measurements of the BAO position. We also fit each mock individually and study the population of BAO fits in order to validate BAO uncertainties and test our method for estimating the covariance matrix of the Lyα forest correlation functions. Finally, we discuss the implications of our results and identify the needs for the next generation of Lyα forest synthetic data sets, with the top priority being to simulate the effect of BAO broadening due to non-linear evolution.

79 ASTRONOMY AND ASTROPHYSICS↗

Translation and rotation equivariant normalizing flow (TRENF) for optimal cosmological analysis

ABSTRACT Our Universe is homogeneous and isotropic, and its perturbations obey translation and rotation symmetry. In this work, we develop translation and rotation equivariant normalizing flow (TRENF), a generative normalizing flow (NF) model which explicitly incorporates these symmetries, defining the data likelihood via a sequence of Fourier space-based convolutions and pixel-wise non-linear transforms. TRENF gives direct access to the high dimensional data likelihood p(x|y) as a function of the labels y, such as cosmological parameters. In contrast to traditional analyses based on summary statistics, the NF approach has no loss of information since it preserves the full dimensionality of the data. On Gaussian random fields, the TRENF likelihood agrees well with the analytical expression and saturates the Fisher information content in the labels y. On non-linear cosmological overdensity fields from N-body simulations, TRENF leads to significant improvements in constraining power over the standard power spectrum summary statistic. TRENF is also a generative model of the data, and we show that TRENF samples agree well with the N-body simulations it trained on, and that the inverse mapping of the data agrees well with a Gaussian white noise both visually and on various summary statistics: when this is perfectly achieved the resulting p(x|y) likelihood analysis becomes optimal. Finally, we develop a generalization of this model that can handle effects that break the symmetry of the data, such as the survey mask, which enables likelihood analysis on data without periodic boundaries.

79 ASTRONOMY AND ASTROPHYSICS↗

The 3D Lyman- α forest power spectrum from eBOSS DR16

We measure the three-dimensional power spectrum (P3D) of the transmitted flux in the Lyman-α (Ly α) forest using the complete extended Baryon Oscillation Spectroscopic Survey data release 16 (eBOSS DR16). This sample consists of ~205 000 quasar spectra in the redshift range 2 ≤ z ≤ 4 at an effective redshift z = 2.334. We propose a pair-count spectral estimator in configuration space, weighting each pair by exp( i k ∙ r), for wave vector k and pixel pair separation r, effectively measuring the anisotropic power spectrum without the need for fast Fourier transforms. This accounts for the window matrix in a tractable way, avoiding artefacts found in Fourier-transform based power spectrum estimators due to the sparse sampling transverse to the line of sight of Ly α skewers. We extensively test our pipeline on two sets of mocks: (i) idealized Gaussian random fields with a sparse sampling of Ly α skewers, and (ii) log-normal LyaCoLoRe mocks including realistic noise levels, the eBOSS survey geometry and contaminants. On eBOSS DR16 data, the Kaiser formula with a non-linear correction term obtained from hydrodynamic simulations yields a good fit to the power spectrum data in the range $(0.02 ≤ k ≤ 0.35)$ h Mpc -1 at the 1–2σ level with a covariance matrix derived from LyaCoLoRe mocks. We demonstrate a promising new approach for full-shape cosmological analyses of Ly α forest data from cosmological surveys such as eBOSS, the currently observing Dark Energy Spectroscopic Instrument and future surveys such as the Prime Focus Spectrograph, WEAVE-QSO, and 4MOST.

79 ASTRONOMY AND ASTROPHYSICS↗

Maximum a posteriori Ly α estimator (MAPLE): band power and covariance estimation of the 3D Ly α forest power spectrum

We present a novel maximum a posteriori estimator to jointly estimate band powers and the covariance of the three-dimensional power spectrum (P3D) of Ly $\alpha$ forest flux fluctuations, called MAPLE. Our Wiener-filter based algorithm reconstructs a window-deconvolved P3D in the presence of complex survey geometries typical for Ly $\alpha$ surveys that are sparsely sampled transverse to and densely sampled along the line of sight. We demonstrate our method on idealized Gaussian random fields with two selection functions: (i) a sparse sampling of 30 background sources per square degree designed to emulate the current Dark Energy Spectroscopic Instrument; (ii) a dense sampling of 900 background sources per square degree emulating the upcoming Prime Focus Spectrograph Galaxy Evolution Survey. Our proof-of-principle shows promise, especially since the algorithm can be extended to marginalize jointly over nuisance parameters and contaminants, i.e. offsets introduced by continuum fitting. Our code is implemented in JAX and is publicly available on GitHub.

79 ASTRONOMY AND ASTROPHYSICS↗