Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Label Space”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Modular flow in JT gravity and entanglement wedge reconstruction

It has been shown in recent works that JT gravity with matter with two boundaries has a type II ∞ algebra on each side. As the bulk spacetime between the two boundaries fluctuates in quantum nature, we can only define the entanglement wedge for each side in a pure algebraic sense. As we take the semiclassical limit, we will have a fixed long wormhole spacetime for a generic partially entangled thermal state (PETS), which is prepared by inserting heavy operators on the Euclidean path integral. Under this limit, with appropriate assumptions of the matter theory, geometric notions of the causal wedge and entanglement wedge emerge in this background. In particular, the causal wedge is manifestly nested in the entanglement wedge. Different PETS are orthogonal to each other, and thus the Hilbert space has a direct sum structure over sub-Hilbert spaces labeled by different Euclidean geometries. The full algebra for both sides is decomposed accordingly. From the algebra viewpoint, the causal wedge is dual to an emergent type III 1 subalgebra, which is generated by boundary light operators. To reconstruct the entanglement wedge, we consider the modular flow in a generic PETS for each boundary. We show that the modular flow acts locally and is the boost transformation around the global RT surface in the semiclassical limit. It follows that we can extend the causal wedge algebra to a larger type III 1 algebra corresponding to the entanglement wedge. Within each sub-Hilbert space, the original type II ∞ reduces to type III 1 .

2D Gravity↗

Morphology-Based Building Use-Type Modeling: Learnability-First Schema Discovery

This technical memorandum documents an update to the building use-type classification workflow, used in LandScan Mosaic, that replaces a fixed, semantically defined class schema with a learnability-first schema discovery procedure. Historically, the target label schema was specified a priori (e.g., predicting a chosen set of use-type codes), and model training and evaluation were performed within that fixed label space. In the updated workflow, the pipeline first evaluates which non-residential distinctions are learnable under spatial generalization and then collapses ambiguous classes into data-driven groupings before finalizing the schema used for production training.

97 MATHEMATICS AND COMPUTING↗

Separable physics-informed DeepONet: Breaking the curse of dimensionality in physics-informed machine learning

The deep operator network (DeepONet) has shown remarkable potential in solving partial differential equations (PDEs) by mapping between infinite-dimensional function spaces using labeled datasets. However, in scenarios lacking labeled data, the physics-informed DeepONet (PI-DeepONet) approach, which utilizes the residual loss of the governing PDE to optimize the network parameters, faces significant computational challenges, particularly due to the curse of dimensionality. This limitation has hindered its application to high-dimensional problems, making even standard 3D spatial with 1D temporal problems computationally prohibitive. Additionally, the computational requirement increases exponentially with the discretization density of the domain. Here, to address these challenges and enhance scalability for high-dimensional PDEs, we introduce the Separable physics-informed DeepONet (Sep-PI-DeepONet). This framework employs a factorization technique, utilizing sub-networks for individual one-dimensional coordinates, thereby reducing the number of forward passes and the size of the Jacobian matrix required for gradient computations. By incorporating forward-mode automatic differentiation (AD), we further optimize computational efficiency, achieving linear scaling of computational cost with discretization density and dimensionality, making our approach highly suitable for high-dimensional PDEs. We demonstrate the effectiveness of Sep-PI-DeepONet through three benchmark PDE models: the viscous Burgers’ equation, Biot’s consolidation theory, and a parameterized heat equation. Our framework maintains accuracy comparable to the conventional PI-DeepONet while reducing training time by two orders of magnitude. Notably, for the heat equation solved as a 4D problem, the conventional PI-DeepONet was computationally infeasible (estimated 289.35 h), while the Sep-PI-DeepONet completed training in just 2.5 h. These results underscore the potential of Sep-PI-DeepONet in efficiently solving complex, high-dimensional PDEs, marking a significant advancement in physics-informed machine learning.

Neural operator↗

Automated Shift Detection in Sensor-Based PV Power and Irradiance Time Series: Preprint

PV power and irradiance sensor-based measurements are prone to error, resulting in issues such as abrupt time series data shifts. These shifts, which are usually unintentional, may be caused by software or hardware configuration changes on a PV system, and do not reflect an actual change in overall system performance. Locating these shifts and segmenting the associated time series aids in more accurate future PV analysis. In this research, an offline changepoint detection (CPD) algorithm that automatically detects these abrupt data shifts in sensor-based time series is introduced. Data shift periods in 101 daily PV power and irradiance time series were labeled manually by two solar experts. These data streams represent sensor-based measurements, and display a variety of data shift behaviors. A changepoint detection algorithm was tuned using the 101 labeled data streams, with each model configuration's ability to detect labeled changepoints benchmarked using metrics such as F1-score, recall, and Rand Index. Best performing models on seasonality-corrected data streams include the Pruned Exact Linear (PELT) method, the Binary Segmentation method, and the Bottom-Up method, all scoring an average F1-score of 0.76 or greater at detecting labeled changepoints within a 30-day window for the labeled data sets. To promote further research in this space, we are releasing the labeled data shift sets on U.S. Department of Energy's (DOE) DuraMAT Data Hub, and the associated algorithm in the Python PVAnalytics package.

changepoint detection↗

Automated Shift Detection in Sensor-Based PV Power and Irradiance Time Series

PV power and irradiance sensor-based measurements are prone to error, resulting in issues such as abrupt time series data shifts. These shifts, which are usually unintentional, may be caused by software or hardware configuration changes on a PV system, and do not reflect an actual change in overall system performance. Locating these shifts and segmenting the associated time series aids in more accurate future PV analysis. In this research, an offline changepoint detection (CPD) algorithm that automatically detects these abrupt data shifts in sensor-based time series is introduced. Data shift periods in 101 daily PV power and irradiance time series were labeled manually by two solar experts. These data streams represent sensor-based measurements, and display a variety of data shift behaviors. A changepoint detection algorithm was tuned using the 101 labeled data streams, with each model configuration's ability to detect labeled changepoints benchmarked using metrics such as F1-score, recall, and Rand Index. Best performing models on seasonality-corrected data streams include the Pruned Exact Linear (PELT) method, the Binary Segmentation method, and the Bottom-Up method, all scoring an average F1-score of 0.76 or greater at detecting labeled changepoints within a 30-day window for the labeled data sets. To promote further research in this space, we are releasing the labeled data shift sets on U.S. Department of Energy's (DOE) DuraMAT Data Hub, and the associated algorithm in the Python PVAnalytics package.

changepoint detection↗

Predicting U 3 O 8 powder processing conditions: An AI/ML approach analyzing deep learning embeddings of SEM micrographs

High-resolution SEM images of uranium-oxide powders encode micro- and nanoscale clues to their synthesis route and calcination temperature. We trained a ResNet-50 model on 11 commercial-scale U₃O₈ classes, ammonium diuranate (ADU) or uranyl peroxide (H₂O₂) precursors calcined at temperatures ranging from 400 to 750 °C and added a 256-D projection head before the classifier to analyze the learned representation. The best of eight seeds reached 92.4 % accuracy on reserved testing data, but our focus is the structure of the embedding space rather than the accuracy and labels. We quantify class relatedness in the original 256-D space using centroid similarity and distributional distances, and we use Uniform Manifold Approximation Projection (UMAP) for visualization. ‘Unknown’ images from different preparation methods, SEM operators, and from the literature localized near the expected classes under a nearest-centroid analysis without retraining, as well as clustered in similar UMAP space. In conclusion, this embedding-centered workflow complements black-box classification by providing quantitative, similarity-based comparisons of U₃O₈ morphologies and reduces storage space by up to 98 % for image data used in millisecond vector search comparisons.

36 MATERIALS SCIENCE↗

A Semi-Supervised Learning Method for the Identification of Bad Exposures in Large Imaging Surveys

As the data volume of astronomical imaging surveys rapidly increases, traditional methods for image anomaly detection, such as visual inspection by human experts, are becoming impractical. We introduce a machine-learning-based approach to detect poor-quality exposures in large imaging surveys, with a focus on the DECam Legacy Survey (DECaLS) in regions of low extinction (i.e., E ( B − V ) < 0.04 ). Our semi-supervised pipeline integrates a vision transformer (ViT), trained via self-supervised learning (SSL), with a k-Nearest Neighbor (kNN) classifier. We train and validate our pipeline using a small set of labeled exposures observed by surveys with the Dark Energy Camera (DECam). A clustering-space analysis of where our pipeline places images labeled in good and bad categories suggests that our approach can efficiently and accurately determine the quality of exposures. Applied to new imaging being reduced for DECaLS Data Release 11, our pipeline identifies 780 problematic exposures, which we subsequently verify through visual inspection. Being highly efficient and adaptable, our method offers a scalable solution for quality control in other large imaging surveys.

Luo, Yufeng (ORCID:0000000246230683)↗

Comparative Assessment of U-Net-Based Deep Learning Models for Segmenting Microfractures and Pore Spaces in Digital Rocks

Segmentation of high-resolution X-ray microcomputed tomography (µCT) images is crucial in digital rock physics (DRP), affecting the characterization and analysis of microscale phenomena in the porous media. The complexity of geological structures and nonideal scanning conditions pose significant challenges to conventional image segmentation approaches. Motivated by the recent increasing popularity of deep learning (DL) techniques in image processing, this work undertakes a comparative study of DL models, specifically U-Net and its variants, for segmenting multiple targets with distinguished features in digital rocks, including discrete fracture networks (DFNs), pore spaces, and solid rock. Particularly, DFNs have a smaller volumetric fraction over others, bringing in a substantial challenge of imbalanced segmentation. The primary focus is to evaluate the architecture and feature enhancement strategies of various DL models, including U-Net, attention U-Net, residual U-Net, U-Net++, and residual U-Net++. The models were designed as 2.5D, utilizing a central 2D image and its two adjacent upper and lower 2D images as input to provide a pseudo-3D context. In addition, because the ground truth of segmentation was unknown for real-world digital rocks, we created a benchmark data set following the inverse operations of segmentation. The data synthesis started from the label images (i.e., solid rock, pore spaces, and DFNs), followed by simulating partial volume blurring, adding random background noise, and introducing ring artifacts to mimic real raw X-ray µCT images. The data set, which included various rock types (i.e., sandstone and artificial data), scanning resolution, and magnitudes of noise and artifacts, was divided into training and testing data sets with a 90% and 10% ratio, respectively. Moreover, in addition to the conventional pixel-wise evaluation metrics, the physics-based metric of the lattice-Boltzmann method (LBM) simulated permeability provided more comprehensive assessments. The results demonstrated that the residual connections, nested architectures, and redesigned skip connections contribute to the model performance and give the residual U-Net++ the highest accuracy. The improvements were mainly on the boundaries and small targets, especially the DFNs, which dominate the interconnectivity and therefore affect the permeability greatly. This study also rigorously evaluated the efficiency and generalization of each model, demonstrating that the sophisticated architectures achieved excellent practicability and maintained robust performance on completely unseen data, ensuring their suitability for diverse and challenging DRP applications.

58 GEOSCIENCES↗

Labels as a feature: Network homophily for systematically annotating human GPCR drug-target interactions

Machine learning has revolutionized drug discovery by enabling the exploration of vast, uncharted chemical spaces essential for discovering novel patentable drugs. Despite the critical role of human G protein-coupled receptors in FDA-approved drugs, exhaustive in-distribution drug-target interaction testing across all pairs of human G protein-coupled receptors and known drugs is rare due to significant economic and technical challenges. This often leaves off-target effects unexplored, which poses a considerable risk to drug safety. In contrast to the traditional focus on out-of-distribution exploration (drug discovery), we introduce a neighborhood-to-prediction model termed Chemical Space Neural Networks that leverages network homophily and training-free graph neural networks with labels as features. We show that Chemical Space Neural Networks’ ability to make accurate predictions strongly correlates with network homophily. Thus, labels as features strongly increase a machine learning model’s capacity to enhance in-distribution prediction accuracy, which we show by integrating labeled data during inference. We validate these advancements in a high-throughput yeast biosensing system (3773 drug-target interactions, 539 compounds, 7 human G protein-coupled receptors) to discover novel drug-target interactions for FDA-approved drugs and to expand the general understanding of how to build reliable predictors to guide experimental verification.

Hansson, Frederik G↗

Monomer-dimer tensor-network basis for qubit-regularized lattice gauge theories

Traditional SU⁡(𝑁) lattice gauge theories (LGTs) can be formulated using an orthonormal basis constructed from the irreducible representations (irreps) 𝑉 𝜆 of the SU⁡(𝑁) gauge symmetry. On a lattice, the elements of this basis are tensor networks comprising dimer tensors on the links labeled by a set of irreps {𝜆 ℓ } and monomer tensors on sites labeled by {𝜆 𝑠 }. These tensors naturally define a local site Hilbert space, ℋ$^𝑔_𝑠$, on which gauge transformations act. Gauss’s law introduces an additional index 𝛼 𝑠 =1,2,…,𝒟⁡(ℋ$^𝑔_𝑠$) that labels an orthonormal basis of the gauge-invariant subspace of ℋ$^𝑔_𝑠$. This monomer-dimer tensor-network (MDTN) basis, |{𝜆 𝑠 },{𝜆 ℓ },{𝛼 𝑠 }⟩, of the physical Hilbert space enables the construction of new qubit-regularized SU⁡(𝑁) gauge theories that are free of sign problems while preserving key features of traditional LGTs. Here, we investigate finite-temperature confinement-deconfinement transitions in a simple qubit-regularized SU(2) and SU(3) gauge theory in 𝑑 =2 and 𝑑 =3 spatial dimensions, formulated using the MDTN basis, and show that they reproduce the universal results of traditional LGTs at these transitions. Additionally, in 𝑑 =1, we demonstrate using a plaquette chain that the string tension at zero temperature can be continuously tuned to zero by adjusting a model parameter that plays the role of the gauge coupling in traditional LGTs.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Nuclear data for space exploration

Understanding the harmful effects of galactic cosmic rays (GCRs) on space exploration requires a substantial amount of nuclear data. Specifically, the interaction of energetic GCR charged particles with spacecraft materials generates secondary radiations that, through energy deposition, can harm astronauts and electronic systems. By identifying the gaps in our knowledge of the relevant nuclear data—such as interaction cross sections—and identifying ways to fill those gaps—with measurements, compilations, evaluations, dissemination, reaction modeling, sensitivity studies, and uncertainty quantification—the safety and viability of space exploration can be improved. This work surveys the state of the art in this interdisciplinary field and identifies promising collaborative research topics that have significant potential to advance our understanding of the effects of the space radiation environment on space exploration.

42 ENGINEERING↗

Frictionless knowledge injection for few-shot learning

Cutting-edge machine learning methods often require large volumes of curated training data, precluding their use in national security problems with rare events in massive datasets. We present a method for incorporating abstract knowledge into models tailored for sparse data. A subject matter expert defines salient concepts using data examples, which are encoded in the model’s embedding space. Models are then trained to respect these concepts. This method enables knowledge injection, yielding effective models with limited labeled data and the ability to assess model sensitivity for subject matter expertise across the nonproliferation mission space, as demonstrated with Raman spectra analysis.

Stomps, Jordan [ORNL] (ORCID:0000000178114479)↗

High-Dimensional Bayesian Optimization via Semi-Supervised Learning with Optimized Unlabeled Data Sampling

We introduce a novel semi-supervised learning approach, named Teacher-Student Bayesian Optimization (TSBO ), integrating the teacher-student paradigm into BO to minimize expensive labeled data queries for the first time. TSBO incorporates a teacher model, an unlabeled data sampler, and a student model. The student is trained on unlabeled data locations generated by the sampler, with pseudo labels predicted by the teacher. The interplay between these three components implements a unique selective regularization to the teacher in the form of student feedback. This scheme enables the teacher to predict high-quality pseudo labels, enhancing the generalization of the GP surrogate model in the search space. To fully exploit TSBO , we propose two optimized unlabeled data samplers to construct effective student feedback that well aligns with the objective of Bayesian optimization. Furthermore, we quantify and leverage the uncertainty of the teacher-student model for the provision of reliable feedback to the teacher in the presence of risky pseudo-label predictions. TSBO demonstrates significantly improved sample-efficiency in several global optimization tasks under tight labeled data budgets. The implementation is available at https://github.com/reminiscenty/TSBO-Official.

Yin, Yuxuan↗

Machine learning-assisted upscaling analysis of reservoir rock core properties based on micro-computed tomography imagery

Optimum solutions for geologic modeling and reservoir simulation in industries such as oil and gas recovery and carbon capture and storage require accurate characterization of reservoir properties, which are often heterogeneous. In this study, high-quality micro-computed tomography (CT) images (1.475-μm/pixel resolution) of a sandstone core acquired from the Bell Creek oil field, USA, were used to provide nondestructive analysis of pore- and core-scale heterogeneity across measurement scales of 94–566 μm. In addition to characterizing the as-received sample, the core sample was flooded with brine to evaluate the capacity of the core sample to receive injected fluids. The micro-CT images were systematically segmented into pore spaces and grains via machine learning (ML) steps including image preprocessing, label creation using a traditional ML method based on limited manual image annotation, and finally U-Net segmentation. The segmented image stacks were reconstructed into digital cubes of various scales of voxel lengths. The 3D porosity values were calculated for all the digital cubes, and the fractal dimensions of the cubes were estimated using a box-counting method. The results showed that smaller cubes had greater heterogeneity and that the porosity values could be accurately estimated by fractal dimension and voxel lengths using ML models. For the core sample with brine flooding, the ratio of pores filled by brine to the total pore space was related to the porosity and could also be accurately estimated by porosity, fractal dimension, and voxel lengths using ML models. In conclusion, the results of this study demonstrate that the concept of fractal dimension can be a useful vector to perform upscaling analysis of sandstone rock heterogeneity from the pore to core scale and that fractal dimensions can be used to estimate porosity values and pore space-filling capacity across those scales.

58 GEOSCIENCES↗

Causal diamonds, cluster polytopes and scattering amplitudes

The “amplituhedron” for tree-level scattering amplitudes in the bi-adjoint φ 3 theory is given by the ABHY associahedron in kinematic space, which has been generalized to give a realization for all finite-type cluster algebra polytopes, labelled by Dynkin diagrams. In this letter we identify a simple physical origin for these polytopes, associated with an interesting (1 + 1)-dimensional causal structure in kinematic space, along with solutions to the wave equation in this kinematic “spacetime” with a natural positivity property. The notion of time evolution in this kinematic spacetime can be abstracted away to a certain “walk”, associated with any acyclic quiver, remarkably yielding a finite cluster polytope for the case of Dynkin quivers. The A n–3 , B n–1 /C n–1 and D n polytopes are the amplituhedra for n-point tree amplitudes, one-loop tadpole diagrams, and full integrand of one-loop amplitudes. We also introduce a polytope D¯ n , which chops the D n polytope in half along a symmetry plane, capturing one-loop amplitudes in a more efficient way.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Photon topology

The topology of photons in vacuum is interesting because there are no photons with k = 0, creating a hole in momentum space. We show that while the set of all photons forms a trivial vector bundle $γ$ over this momentum space, the R and L photons form topologically nontrivial subbundles $γ±$ with first Chern numbers ∓2. In contrast, $γ$ has no linearly polarized subbundles, and there is no Chern number associated with linear polarizations. It is a known difficulty that the standard version of Wigner’s little group method produces singular representations of the Poincaré group for massless particles. By considering representations of the Poincaré group on vector bundles we obtain a version of Wigner’s little group method for massless particles which avoids these singularities. Here we show that any massless bundle representation of the Poincaré group can be canonically decomposed into irreducible bundle representations labeled by helicity, which in turn can be associated to smooth irreducible Hilbert space representations. This proves that the R and L photons are globally well defined as particles and that the photon wave function can be uniquely split into R and L components. This formalism offers a method of quantizing the electromagnetic field without invoking discontinuous polarization vectors as in the traditional scheme. We also demonstrate that the spin-Chern number of photons is not a purely topological quantity. Lastly, there has been an extended debate on whether photon angular momentum can be split into spin and orbital parts. Our work explains the precise issues that prevent this splitting. Photons do not admit a spin operator; instead, the angular momentum associated with photons’ internal degree of freedom is described by a helicity-induced subalgebra corresponding to the translational symmetry of $γ$.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Contrastive Machine Learning with Gamma Spectroscopy Data Augmentations for Detecting Shielded Radiological Material Transfers

Data analysis techniques can be powerful tools for rapidly analyzing data and extracting information that can be used in a latent space for categorizing observations between classes of data. Machine learning models that exploit learned data relationships can address a variety of nuclear nonproliferation challenges like the detection and tracking of shielded radiological material transfers. The high resource cost of manually labeling radiation spectra is a hindrance to the rapid analysis of data collected from persistent monitoring and to the adoption of supervised machine learning methods that require large volumes of curated training data. Instead, contrastive self-supervised learning on unlabeled spectra can enhance models that are built on limited labeled radiation datasets. This work demonstrates that contrastive machine learning is an effective technique for leveraging unlabeled data in detecting and characterizing nuclear material transfers demonstrated on radiation measurements collected at an Oak Ridge National Laboratory testbed, where sodium iodide detectors measure gamma radiation emitted by material transfers between the High Flux Isotope Reactor and the Radiochemical Engineering Development Center. Label-invariant data augmentations tailored for gamma radiation detection physics are used on unlabeled spectra to contrastively train an encoder, learning a complex, embedded state space with self-supervision. A linear classifier is then trained on a limited set of labeled data to distinguish transfer spectra between byproducts and tracked nuclear material using representations from the contrastively trained encoder. The optimized hyperparameter model achieves a balanced accuracy score of 80.30%. Any given model—that is, a trained encoder and classifier—shows preferential treatment for specific subclasses of transfer types. Regardless of the classifier complexity, a supervised classifier using contrastively trained representations achieves higher accuracy than using spectra when trained and tested on limited labeled data.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗