Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “embedding model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Self-supervised and multi-fidelity learning for extended predictive soil spectroscopy

Infrared spectroscopy is a cost-effective, non-destructive, and environmentally benign technology that is increasingly recognized as an important solution for meeting the global demand for soil data. While both near-infrared (NIR) and mid-infrared (MIR) diffuse reflectance spectroscopy enable rapid estimation of soil properties, they present a significant trade-off: NIR offers superior scalability and lower operational costs, whereas MIR provides higher analytical fidelity by capturing fundamental molecular vibrations. In this study, we propose a self-supervised, multi-fidelity learning framework designed to bridge this gap. Our approach leverages large-scale MIR spectral libraries to learn a compact, transferable latent representation, into which NIR spectra are subsequently aligned for downstream prediction. The workflow consists of pretraining a latent model on a large MIR library, adapting the representation using a smaller paired NIR–MIR dataset, and evaluating generalization on an independent external test set. Across a range of chemical and physical soil properties, we found that MIR-derived embeddings improved prediction accuracy relative to baseline models that used raw MIR inputs. Predictions derived from the spectrum conversion (NIR to MIR) task did not match the performance of the original MIR spectra but were similar or superior to predictive performance of NIR-only models, suggesting the unified spectral latent space can effectively leverage the larger and more diverse MIR dataset for prediction of soil properties not well represented in current NIR libraries.

54 ENVIRONMENTAL SCIENCES↗

Analytic Neural Network Gaussian Process Enabled Chance-Constrained Voltage Regulation for Active Distribution Systems with PVs, Batteries and EVs

This paper proposes an analytic neural network Gaussian process (NNGP)-based chance-constrained real-time voltage regulation method for active distribution systems with photovoltaics (PVs), batteries, and electric vehicles (EVs). NNGP can utilize historical measurement data to achieve real-time probabilistic node voltage estimation through Bayesian inference. Then, NNGP is fully analytically embedded into the optimal power flow model to perform voltage regulation and adapt to various topological changes. The uncertainties of voltage estimations are easily considered via the chance constraint, and it has been shown that the adoption of this chance constraint can significantly improve the reliability of voltage regulation under various scenarios. The comparison results with other methods, carried out on a real 759-node distribution system located in western Colorado, U.S., show that the proposed method can achieve accurate voltage estimation across different topologies and reliably perform voltage regulation considering PVs, batteries, and EVs.

active distribution systems↗

Detecting Living-off-the-land Attacks Using K-means And Graph Convolutional Networks

The code ingests Zeek logs derived from network packet captures and goes through data preprocessing before it gets passed into a K-Means model that labels each device as either a client or server. Graph Convolutional Network (GCN) model is used to obtain the embeddings to represent the features in lower dimension. Last, K-means cluster analysis is used to cluster the embeddings for each class.

Quach, Anna [Idaho National Laboratory (INL), Idah↗

Flux REaction TArget Prioritization (Flux RETAP) v1

Metabolic engineering is evolving rapidly as a result of new advances in synthetic biology and automation, as well as the irruption of machine learning (ML). ML has been shown to provide the predictive power synthetic biology lacked and needed, and to be able to effectively guide the metabolic engineering process. However, current technical limitations prevent the independent application of ML approaches to metabolic engineering without the use of previous biological knowledge in the form of a prioritized list of desirable engineering targets. Here, we present FluxRETAP, a simple and computationally inexpensive method that leverages the prior mechanistic knowledge embedded in genome-scale metabolic models (GSMs) for suggesting targets for genetic overexpression, downregulation or deletion, with the final goal of increasing metabolite production. FluxRETAP captured 100% of reaction targets experimentally verified to improve Escherichia coli isoprenol production in the literature accessible to us, 50% of targets that experimentally improved taxadiene production in E. coli and ~60% of genetic targets from a verified minimal constrained cut-set in Pseudomonas putida while providing additional high priority targets that could be tested. Overall, FluxRETAP is an efficient algorithm for identifying a prioritized list of testable genetic and reaction targets which can also be utilized in ML pipelines.

Czajka, Jeffrey [Battelle Memorial Institute, Paci↗

Life Cycle Greenhouse Gas Emissions of Coal-Biomass Co-Firing Power Plants with Carbon Capture and Storage

The United States has set a target to achieve the net-zero economy by 2050. Bioenergy with Carbon Capture and Sequestration (BECCS) is one of the promising negative-emission routes in the mitigation portfolio to help meet this goal. Coal-biomass co-firing with carbon capture and storage (CCS) is a key BECCS technology to realize the carbon mitigation at fossil-fuel power plants. The mitigation potential of co-firing option is affected by numerous critical factors, such as biomass properties, co-firing level, and carbon capture rate. The objectives of the study are to characterize and estimate the life cycle greenhouse gas (GHG) emissions and performance of coal-biomass co-firing power plants with CCS, determine the breakeven co-firing level at power plants necessary to achieve net-zero life cycle emissions, and quantify the variabilities and uncertainties in life cycle emissions. The scope of the life cycle assessment includes the fuel supply, combustion-based power generation, and CO2 transport and storage. A fuel-based life cycle module is developed and embedded in the Integrated Environmental Control Model (IECM), a fossil-fuel power plant modeling tool. This study then applies the enhanced IECM to conduct the process-based life cycle assessment for an array of biomass co-firing scenarios. Deterministic analysis indicates that reaching net-zero life cycle emissions in a biomass co-firing plant without CCS deployment is challenging. Combining biomass co-firing and CCS deployment can significantly lower the overall life cycle emissions of power plants. Net-zero life cycle emissions can be achieved with a 20 wt.% co-firing level and 90% CCS when the Powder River Basin coal is co-fired with energy crops or forestry residues. However, the breakeven co-firing level for net-zero emissions depend on the selected fuel properties. Fuel supply and plant operation are the critical stages influencing the life cycle emissions of power plants with 90% CCS. Deployment of deep CCS beyond 90% CO2 capture can remarkably reduce operational emissions and the breakeven co-firing level. With 99% CCS, the breakeven co-firing rate can be reduced to 12% on average. These findings highlight the trade-offs between technical performance and environmental impact of biomass co-firing at coal-fired power plants and emphasize the role of deep CCS in achieving a net-zero emissions future.

Wu, Wanying↗

Shakeup and shakeoff spectra in the electron capture decay of atomic 7 Be

The most stringent laboratory-based experimental limits on the existence of submegaelectronvolt sterile neutrinos are currently set by decay spectroscopy of radioactive 7 Be embedded into superconducting sensors. The systematic uncertainties are dominated by the modeling of the electron shakeup and shakeoff spectra that are not based on state-of-the-art atomic theory and do not include electron correlations or relativistic effects. We have used the multiconfiguration Dirac-Fock formalism to obtain correlated wave functions ab initio and compute all single and double shake processes in the electron capture decay of atomic 7 Be . The simulations can explain some but not all of the observed spectral features, likely because the wave functions are modified by the Ta sensor material that the 7 Be is embedded into. The new models also show that the 𝐿/𝐾 electron capture ratio of 7 Be in Ta has previously been slightly underestimated revising the previous value of 0.070(7) to a new value of 0.0756(20).

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Shake-up and shake-off spectra in the electron capture decay of atomic $^7$Be

The most stringent laboratory-based experimental limits on the existence of sub-MeV sterile neutrinos are currently set by decay spectroscopy of radioactive $^7$Be embedded into superconducting sensors. The systematic uncertainties are dominated by the modeling of the electron shake-up and shake-off spectra that are not based on state-of-the-art atomic theory and do not include electron correlations or relativistic effects. We have used the multiconfiguration Dirac-Fock formalism to obtain correlated wavefunctions ab initio and compute all single and double shake processes in the electron capture decay of atomic $^7$Be. The simulations can explain some but not all of the observed spectral features, likely because the wave functions are modified by the Ta sensor material that the $^7$Be is embedded into. The new models also show that the L/K electron capture ratio of $^7$Be in Ta has previously been slightly underestimated revising the previous value of 0.070(7) to a new value of 0.0756(20).

Atomic Physics (physics.atom-ph)↗

Fracture Analysis of Cohesive Zone Models for Modeling Residual Stress Induced Delamination in Composite Structures

A fracture study of coupon-scale composite cylinders with embedded defects was conducted with an objective to assess and validate a modeling approach using two available cohesive material models. The study included experimental and simulation evaluations of initiation of crack growth and progression. Interrupted thermal experiments used acoustic emissions monitoring to identify the onset of crack progression during each cooling interval and ultrasonic scanning provided images of defect growth. Verification, validation, and uncertainty quantification (VVUQ) processes were performed in the assessment of the simulation predicted temperature at which crack propagation begins (quantity of interest). The Sobol sensitivity analysis identified the hoop direction elastic modulus in the carbon fiber reinforced polymer (CFRP) plies as the most influential parameter for simulations using both cohesive models, accounting for at least 70% of the variation in the temperature at crack propagation. The UQ temperature range for the Tvergaard-Hutchinson model was higher (more conservative) than the experimental acoustic measurement indicators of crack progression, while the temperature range for the Thouless-Parmigiani model enveloped the experimental data points for the primary defect size of 0.75 x 1 in. The simulations could not capture the stable crack growth indicated in the experiments. This is likely due to the models’ inability to represent anisotropic fracture toughness attributed to the structure of the orthotropic fiber weave in a woven composite laminate.

42 ENGINEERING↗

Understanding Subseasonal Moisture Recycling from Extreme MCSs Using a Land–Atmosphere Coupled Water Tracer Model

Extreme precipitation events often produce severe flooding and pose significant threats to our society. However, their subseasonal impacts on the terrestrial and atmospheric systems are not well understood. Here, we use a new land–atmosphere coupled water tracer tool embedded in the Weather Research and Forecasting (WRF) Model to “tag” precipitation produced by a series of extreme mesoscale convective systems (MCSs) in May 2015 over the southern Great Plains to reveal their contributions to different terrestrial water storages and moisture recycling during May–July. During May, we find that precipitation from earlier MCSs can contribute to 20%–50% of total evapotranspiration (ET) in later May and contribute ∼10% of local precipitation for later MCSs. After May, the water “tagged” to preceding May MCS events has the most active contribution to moisture recycling during the first 5 days in June, but this contribution diminishes quickly afterward. Both periods indicate a quick turnover of tagged precipitation, which may be representative of the southern Great Plains region. The quick turnover is closely associated with the direct evaporation from the soil surface due to the predominant shrubland and grassland coverage over this region. Over some forested areas, however, the tracer-ET flux is dominated by transpiration that stably supplies a small fraction (<1%) of moisture to the atmosphere throughout June and July. In conclusion, the tracer-contributed plume (through evaporation and transpiration) in atmospheric moisture can extend 500–700 km downstream between May and July, highlighting the far-reaching contribution to moisture recycling and precipitation from the tagged MCS events in the southern Great Plains.

Atmosphere-land interaction↗

Debunking common myths in coastal circulation modeling

Despite tremendous progress in algorithm development, computational efficiency and transition into operations over the past two decades, coastal modeling still lacks scientific rigor due to proliferation of many ‘gray’ areas related to various modeling choices made by modelers. Here, in this paper, we propose some guiding principles for the modeling community to improve performance, and we also debunk commonly held myths that make the coastal modeling lack rigor. Using our own experience in developing seamless cross-scale unstructured-grid based models for the past two decades, we describe in unprecedented detail the end-to-end modeling process (i.e., from digital elevation models (DEMs) to mesh generation to post analysis), and demonstrate that defensible modeling is within reach for any end user by following three guiding principles: (1) Bathymetry is a first order forcing in coastal domains and thus should be respected in all aspects of modeling; (2) Oceanographic processes are driven across multiple spatial scales and so models should enable appropriate resolution as needed; and (3) Model assessment should focus on physical processes. Through qualitative and quantitative model assessments, we demonstrate the fundamental role played by bathymetry/topography as embedded in DEMs in making the results defensible, which is unfortunately glossed over in many modeling studies. Focusing on process-based assessment simplifies the calibration process. A major conclusion of this work is that model developers and operators should maximize the scientific rigor for in silico oceanography by avoiding some common pitfalls that rely on error compensation at the expense of representation of physical system processes. We present some best practice procedures for defensive and trustworthy numerical modeling.

54 ENVIRONMENTAL SCIENCES↗

SetBERT: the deep learning platform for contextualized embeddings and explainable predictions from high-throughput sequencing

MOTIVATION: High-throughput sequencing (HTS) is a modern sequencing technology used to profile microbiomes by sequencing thousands of short genomic fragments from the microorganisms within a given sample. This technology presents a unique opportunity for artificial intelligence to comprehend the underlying functional relationships of microbial communities. However, due to the unstructured nature of HTS data, nearly all computational models are limited to processing DNA sequences individually. This limitation causes them to miss out on key interactions between microorganisms, significantly hindering our understanding of how these interactions influence the microbial communities as a whole. Furthermore, most computational methods rely on post-processing of samples which could inadvertently introduce unintentional protocol-specific bias. RESULTS: Addressing these concerns, we present SetBERT, a robust pre-training methodology for creating generalized deep learning models for processing HTS data to produce contextualized embeddings and be fine-tuned for downstream tasks with explainable predictions. By leveraging sequence interactions, we show that SetBERT significantly outperforms other models in taxonomic classification with genus-level classification accuracy of 95%. Furthermore, we demonstrate that SetBERT is able to accurately explain its predictions autonomously by confirming the biological-relevance of taxa identified by the model. AVAILABILITY AND IMPLEMENTATION: All source code is available at https://github.com/DLii-Research/setbert. SetBERT may be used through the q2-deepdna QIIME 2 plugin whose source code is available at https://github.com/DLii-Research/q2-deepdna.

Ludwig, David W↗

Street-level temperature estimation using graph neural networks: Performance, feature embedding and interpretability

Estimating street-level air temperature is a challenging task due to the highly heterogeneous urban surfaces, canyon-like street morphology, and the diverse physical processes in the built environment. Though pioneering studies have embarked on investigations via data-driven approaches, many questions remain to be answered. Here, in this study, we leveraged an innovative framework and redefined the street-level temperature estimation problem using Graph Neural Networks (GNN) with spatial embedding techniques. The results showed that GNN models are more capable and consistent of estimating street-level temperature among tested locations, benefiting from its unique strength in handling extensive data over unstructured graph topology. In addition, we conducted in-depth analysis of feature importance to enhance the model interpretability. Among the urban features analyzed in this study, the time-variant canopy density and meter-level land use data emerge as crucial factors. Our findings highlight GNN 's high potential in capturing the complex dynamics between urban elements and their impacts on microclimate, thus offering valuable insights for comprehensive urban data collection and urban climate modeling in general. Collectively, this study also contributes to urban planning and policy by providing avenues to enhance city resilience against climate change, thereby advancing the agenda for environmental stewardship and urban sustainability.

54 ENVIRONMENTAL SCIENCES↗

Machine learning for seismic low-frequency extrapolation

The cycle-skipping problem that plagues full waveform inversion (FWI) can be at least partially mitigated if low frequencies (which encode the kinematics of wave propagation in seismic data) are recorded. However, seismic sources and receivers are band-limited, so seismic data does not generally include signals down to 0 Hz. To improve our ability to solve the seismic inverse problem, one can synthesize this missing low-frequency (LF) content from the recorded high-frequency (HF) data using machine learning (ML) models. Deep learning models such as convolutional neural networks (CNNs) demonstrate impressive ability to perform low frequency extrapolation. However, such models require powerful hardware (GPU machines) and careful training. We assess the extrapolation capabilities of three different ML models that do not require GPU machines, namely, random forest, Gaussian process regression and gradient boosting, on both synthetic and real data. Experimental results on two synthetic data sets (generated from a low velocity lens embedded in a homogeneous medium, and the Marmousi model) demonstrate that FWI applied to the extrapolated data consistently improves inversion accuracy relative to FWI applied to the original data sets that do not contain low frequencies. Application of low-frequency extrapolation to real data from the Northwest Shelf of Australia demonstrates that tree-based ML models such as gradient boosting can outperform CNNs in terms of both accuracy and computational cost on non-GPU architectures.

58 GEOSCIENCES↗

Stakeholder-guided holistic, Adaptive Framework for enhancing community Energy Resilience (SAFER) (Final Technical Report)

The Stakeholder-guided holistic, Adaptive Framework for enhancing community Energy Resilience (SAFER) project advances resilience science and engineering by addressing challenges in rural Kansas communities where aging infrastructure, extreme weather, and socioeconomic disparities heighten vulnerability to energy disruptions. Traditional approaches often focus on technical performance while overlooking community concerns and priorities. SAFER responds by integrating community perspectives with advanced analytical frameworks to create a holistic model for measuring and improving resilience. Project objectives included developing novel resilience metrics, advancing modeling frameworks that capture interdependencies across infrastructures, and embedding community-centric indicators directly into planning processes for distributed energy resources. The key technical innovations included the creation of self-organizing map (SOM)-based indices for objective resilience quantification, hetero-functional graph theory (HFGT) models linking power, water, transportation, and community assets, and graph neural network (GNN) tools for identifying critical nodes in complex systems. Community-centric energy planning was demonstrated through optimal siting and sizing of (photovoltaic) PV and battery storage, ensuring resilience enhancements also addressed energy burden and energy insecurity. SAFER engaged community partners in Dodge City and Ford County through surveys, focus groups, and workshops, generating more than 600 responses that established baseline measures of energy burden, financial insecurity, and willingness-to-pay to avoid outages. This data, organized in terms of a community capitals framework, informed the development of weighted reliability indices that better reflect community costs than traditional utility metrics. SAFER’s GNN-based critical node identification framework identified expert-labelled critical nodes with over 99% accuracy, while also uncovering additional functionalities essential for proactive resilience planning. The project’s models demonstrated that optimal PV and storage deployment could improve resilience indices by over 11 percent, with dispatch strategies further enhancing outcomes, confirming both the technical effectiveness and economic feasibility of these approaches. Through its combined emphasis on rigorous modeling, community-focused planning, and community engagement, SAFER advances the state of resilience research while delivering direct benefits to rural communities. The project provides tools, guidelines, and resilience heatmaps that help utilities, local governments, and residents better anticipate disruptions, prioritize investments, and strengthen the capacity to withstand and recover from energy-related hazards. Furthermore, the developed HFG and GNN frameworks are designed for transferability, allowing them to be adapted for resilience planning in other communities with minimal retraining. This inductive learning capability provides a scalable pathway to extend the SAFER project’s impact. Thus, creating a foundation for a nationally applicable model of infrastructure resilience. Additionally, the HFG can also be extended to include other FEMA community lifelines.

14 SOLAR ENERGY↗

Optical properties of a diamond NV color center from capped embedded multiconfigurational correlated wavefunction theory

Diamond defects are among the most promising qubits. Modeling their properties through accurate quantum mechanical simulations can further their development into robust units of information. We use the recently developed capped density functional embedding theory (capped-DFET) with the multiconfigurational n-electron valence second-order perturbation theory to characterize the electronic excitation energies for different spin manifolds of the well-characterized negatively charged substitutional N defect adjacent to a vacancy (V C ) in diamond (N C V C − ). We successfully reproduce vertical excitation energies for both triplet and singlet states of N C V C − with errors < 0.1 eV. Unlike other embedding methods, capped-DFET exhibits robust predictions that are approximately independent of the embedded cluster size: it only requires a cluster to contain the defect atoms and their nearest neighbors (as small as a 40-atom capped cluster). Furthermore, our method is free from slowly converging Coulomb interactions between charged defects, and thus also only weakly dependent on supercell size.

Chemistry↗

Sequential Kalman tuning of the t -preconditioned Crank-Nicolson algorithm: efficient, adaptive and gradient-free inference for Bayesian inverse problems

Ensemble Kalman Inversion (EKI) has been proposed as an efficient method for the approximate solution of Bayesian inverse problems with expensive forward models. However, when applied to the Bayesian inverse problem EKI is only exact in the regime of Gaussian target measures and linear forward models. Here, in this work we propose embedding EKI and Flow Annealed Kalman Inversion, its normalizing flow (NF) preconditioned variant, within a Bayesian annealing scheme as part of an adaptive implementation of the t-preconditioned Crank-Nicolson (tpCN) sampler. The tpCN sampler differs from standard pCN in that its proposal is reversible with respect to the multivariate t-distribution. The more flexible tail behaviour allows for better adaptation to sampling from non-Gaussian targets. Within our Sequential Kalman Tuning (SKT) adaptation scheme, EKI is used to initialize and precondition the tpCN sampler for each annealed target. The subsequent tpCN iterations ensure particles are correctly distributed according to each annealed target, avoiding the accumulation of errors that would otherwise impact EKI. We demonstrate the performance of SKT for tpCN on three challenging numerical benchmarks, showing significant improvements in the rate of convergence compared to adaptation within standard SMC with importance weighted resampling at each temperature level, and compared to similar adaptive implementations of standard pCN. The SKT scheme applied to tpCN offers an efficient, practical solution for solving the Bayesian inverse problem when gradients of the forward model are not available. Code implementing the SKT schemes for tpCN is available at https://github.com/RichardGrumitt/KalmanMC.

97 MATHEMATICS AND COMPUTING↗

Infrasonic directivity of monopole, dipole and bipole ground-surface reflected sources

Infrasound (acoustic waves below 20 Hz) can be used to detect, locate and quantify activity in the atmosphere such as volcanic eruptions and anthropogenic explosions. Attempts to quantify volcanic eruption parameters such as exit velocity, plume height and mass flow rate using infrasound data depend strongly on assumptions of the acoustic source type. Infrasonic sources may produce omnidirectional or directional wavefields, while propagation effects, such as interaction with topography, can induce further wavefield directivity that is measured by field instrumentation. Limited sampling of these wavefields can hinder our ability to infer the underlying source, and thus our understanding of the eruption characteristics. Equivalent sources are often used to represent acoustic source mechanisms and resultant wavefields. In this study, we review equivalent acoustic sources as they pertain to infrasonic scale and wavelengths commonly encountered in very local (⁠<5 km range) geophysical field deployments. We highlight the equivalent infrasonic bipole source that can be induced by ground-reflection of an elevated monopole; we are not aware of any prior infrasound studies that use the bipole source concept. We use analytical and numerical methods to explore source directivity of monopole, dipole and bipole ground-reflected sources at infrasonic frequencies as well as the additional directivity complications introduced by interactions with topography. We illustrate that for typical volcano-infrasound wavelengths, increasing height above the ground as well as increasing source frequency leads to increased wavefield directivity. Numerical modelling using a simple omnidirectional monopole source embedded in topography further illustrates that both horizontal and vertical infrasound directionality can be induced by topography at the distance scales appropriate for local volcano infrasound monitoring. Information summarized in this analytical and numerical exploration of infrasound directivity may be used to help guide future volcano-infrasound field deployments intended to estimate source parameters or quantify wavefield directivity. Analytic solutions for simple whole-space or half-space atmospheres provide useful formulations for planning or initially analysing geophysical field-scale experimental data; however, especially at very local distances from the source (⁠<5 km), 3-D simulations are necessary to account for complex topography commonly encountered in volcano-infrasound applications.

Infrasound↗