Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Dimensionality reduction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

SO(3)-invariant PCA with application to molecular data

Principal component analysis (PCA) is a fundamental technique for dimensionality reduction and denoising; however, its application to three-dimensional data with arbitrary orientations -- common in structural biology -- presents significant challenges. A naive approach requires augmenting the dataset with many rotated copies of each sample, incurring prohibitive computational costs. In this paper, we extend PCA to 3D volumetric datasets with unknown orientations by developing an efficient and principled framework for SO(3)-invariant PCA that implicitly accounts for all rotations without explicit data augmentation. By exploiting underlying algebraic structure, we demonstrate that the computation involves only the square root of the total number of covariance entries, resulting in a substantial reduction in complexity. We validate the method on real-world molecular datasets, demonstrating its effectiveness and opening up new possibilities for large-scale, high-dimensional reconstruction problems.

Fraiman, Michael [Tel Aviv Univ., Tel Aviv (Israel

A Class of Sparse Johnson–Lindenstrauss Transforms and Analysis of their Extreme Singular Values

The Johnson–Lindenstrauss (JL) lemma is a powerful tool for dimensionality reduction in modern algorithm design. The lemma states that any set of high-dimensional points in a Euclidean space can be projected into lower dimensions while approximately preserving pairwise Euclidean distances. Random matrices satisfying this lemma are called JL transforms (JLTs). Inspired by existing $s$-hashing JLTs with exactly $s$ nonzero elements on each column, the present work introduces an ensemble of sparse matrices encompassing so-called $s$-hashing-like matrices whose expected number of nonzero elements on each column is $s$. The independence of the sub-Gaussian entries of these matrices and the knowledge of their exact distribution play an important role in their analyses. Using properties of independent sub-Gaussian random variables, these matrices are demonstrated to be JLTs, and their smallest nontrivial singular values and largest singular values are estimated nonasymptotically using a technique from geometric functional analysis. As the dimensions of the matrix grow to infinity, these singular values are proved to converge almost surely to fixed quantities (by using the universal Bai–Yin law) and in distribution to the Gaussian orthogonal ensemble Tracy–Widom law after proper rescalings. Understanding the behaviors of extreme singular values is important in general because they are often used to define a measure of stability of matrix algorithms. For example, JLTs were recently used in derivative-free optimization algorithmic frameworks to select random subspaces in which are constructed random models or poll directions to achieve scalability, and hence estimating their smallest singular value in particular helps determine the dimension of these subspaces.

97 MATHEMATICS AND COMPUTING

Uncertainty propagation and sensitivity analysis for constrained optimization of nuclear waste vitrification

Abstract The vitrification of high‐level waste (HLW) by heating a mixture of glass‐forming chemicals (GFCs) with the waste can be improved using a constrained optimization problem. This study explores how different uncertainty propagation (UP) methods implemented with the optimization process can affect the glass formulation of nuclear waste glasses. UP is the effort of propagating uncertain inputs through a system to understand and quantify output distributions. Uncertainty intervals are crafted from output distributions to inform the optimization algorithm. UP is often implemented with Monte Carlo (MC) sampling for large nonlinear systems, which can be difficult to implement within a constrained optimization algorithm that requires derivative information. Other UP methods often used for optimization under uncertainty (OUU) can be designed to work within an established constrained optimization framework. Methods of UP are evaluated in this study including iterative sampling approaches, first‐order approximations, and surrogate modeling with machine learning (ML). A method of dimensional reduction based on global sensitivity analysis is introduced to support the UP methods for the large dimensionality of the problem. Analytical UP methods able to achieve similar optimums 10 times faster than the baseline MC approach, and produce 93.9% similar output distributions are reported.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W

Experimental Realization of One-Dimensional Helium

Abstract As the spatial dimension is lowered, locally stabilizing interactions are reduced, leading to the emergence of strongly fluctuating phases of matter without classical analogues. Realizing 1D platforms has been elusive, due to their inherent lack of stability, with a few notable exceptions such as spin chains and ultracold low-density gasses. The inability of such systems to exhibit long range order is essential to their universal description in terms of the Tomonaga-Luttinger liquid theory. Here we report on the experimental observation of a one-dimensional quantum liquid of $$^4$$ 4 He using nanoengineering to confine it within a porous material preplated with a noble gas to enhance dimensional reduction. The resulting excitations of the confined $$^4$$ 4 He, confirmed by neutron scattering, are qualitatively different than three- and two-dimensional superfluid helium, and consistent with Quantum Monte Carlo calculations. The results can be analyzed in terms of a mobile impurity in an otherwise linear Luttinger liquid allowing for the extraction of the microscopic parameters describing the emergent quantum liquid.

Sokol, Paul E.

In Situ Data Analysis Through Physics-informed Tensor Decompositions (LDRD Final Report)

We introduce a new low-dimensional model of high-dimensional numerical simulation data based on low-rank tensor decompositions. Our new model aims to minimize differences between the model data and simulation data as well as functions of the model data and functions of the simulation data. This novel approach to dimensionality reduction of simulation data provides a means of directly incorporating quantities of interests and invariants associated with conservation principles associated with the simulation data into the low-dimensional model, thus enabling more accurate analysis of the simulation without requiring access to the full set of high-dimensional data. Computational results of applying this approach to two standard low-rank tensor decompositions of data arising from simulation of combustion and plasma physics are presented.

97 MATHEMATICS AND COMPUTING

A two-stage optical fusion framework for wildfire severity mapping across the conterminous United States

Accurate wildfire severity mapping (WSM) is essential for post-fire recovery planning, erosion risk assessment, ecosystem monitoring, and disaster risk reduction. Although Landsat and Sentinel optical imagery have been widely used for burn severity assessment, the added value of fusing multiple optical sensors has not been sufficiently quantified across diverse fire events, particularly since the launch of Landsat-9. This study evaluates whether multisensor optical fusion improves wildfire severity mapping relative to single-sensor baselines using Sentinel-2, Landsat-8, and Landsat-9 imagery across 40 wildfire events in the conterminous United States. We tested a two-stage fusion framework that combines feature-level fusion with pixel-level dimensionality reduction. First, feature-level fused datasets were created through early fusion by combining standardized post-fire bands from each sensor into a single predictor stack. Both raw reflectance bands and pairwise spectral transforms were retained to capture within- and cross-sensor spectral interactions. Second, Linear Discriminant Analysis was applied to both single-sensor and fused datasets to produce comparable low-dimensional feature spaces. Six machine-learning classifiers were then used to benchmark model performance with repeated spatially buffered train–test splits. Results show that Landsat-9 was the strongest single-sensor baseline. Among the fusion strategies, Sentinel-2 + Landsat-9 produced the most consistent improvement and reduced performance variability. Landscape-condition analysis further showed that this fusion was most beneficial in shrubland-dominated and high-terrain fires, where it achieved the highest overall mean accuracy and the fewest failures. In contrast, its benefits were less reliable in evergreen forests, mixed vegetation, and low- to moderate-elevation terrain. In operational settings, the Sentinel-2 + Landsat-9 configuration offers a practical solution for post-fire recovery planning, erosion-risk assessment, watershed management, and ecological monitoring when field observations are available and timely satellite-based information is needed.

Landsat

Layered and Low-Dimensional Lead, Silver, and Bismuth Halide Perovskites Directed by Halogen-Substituted Spacer Cations

Hybrid organic–inorganic metal halides provide a diverse parameter space in which the optoelectronic properties can be tuned through the composition. The compositional tunability extends to the metal site, which can be expanded from single valent metals (e.g., Pb 2+ ) to multivalent metals (e.g., Ag + and Bi 3+ ), and the dimension (2D, 1D, or 0D). However, a deeper understanding of how the organic cations template these metal halide structures is needed. Here, we synthesize and study the structures of a series of new layered and low-dimensional metal (Pb, Ag, and Bi) halides templated by the halogenated aryl spacer cations 2-chlorobenzylammonium (2ClBZ) and 3-chloro-2-fluorobenzylammonium (3Cl2FBZ). We report new lead perovskites, (3Cl2FBZ) 2 PbBr 4 , (2ClBZ) 3 PbI 5 , and (3Cl2FBZ) 2 PbI 4 , and compare them to their silver and/or bismuth analogs (2ClBZ) 4 AgBiBr 8 , (3Cl2FBZ) 4 AgBiBr 8 , (2ClBZ) 3 Bi 2 I 9 , and (3Cl2FBZ) 4 Bi 2 I 10 . In all structures, the halogen-substituted cations result in 2D or “pseudo-2D” layering, but the different halogen substituents introduce different distortions (tilting, octahedral distortion) and dimensional reduction to 1D or 0D depending on the metal and halide compositions. Optical absorption measurements reveal the bandgaps are tunable through metal sites, dimension, and cations to different extents. Furthermore, the 1D (3Cl2FBZ) 4 Bi 2 I 10 crystallizes in the noncentrosymmetric space group Cmc2 1 and exhibits second-harmonic generation (SHG). Furthermore, the organic–inorganic interactions and resultant structural distortions examined here provide insights toward the engineering of noncentrosymmetry and dimensional control in hybrid metal halide perovskites.

Cations

Machine-learned closure of URANS for stably stratified turbulence: connecting physical timescales & data hyperparameters of deep time-series models

Stably stratified turbulence (SST), a model that is representative of the turbulence found in the oceans and atmosphere, is strongly affected by fine balances between forces and becomes more anisotropic in time for decaying scenarios. Moreover, there is a limited understanding of the physical phenomena described by some of the terms in the Unsteady Reynolds-Averaged Navier–Stokes (URANS) equations—used to numerically simulate approximate solutions for such turbulent flows. Rather than attempting to model each term in URANS separately, it is attractive to explore the capability of machine learning (ML) to model groups of terms, i.e. to directly model the force balances. We develop deep time-series ML for closure modeling of the URANS equations applied to SST. We consider decaying SST which are homogeneous and stably stratified by a uniform density gradient, enabling dimensionality reduction. We consider two time-series ML models: long short-term memory and neural ordinary differential equation. Both models perform accurately and are numerically stable in a posteriori (online) tests. Furthermore, we explore the data requirements of the time-series ML models by extracting physically relevant timescales of the complex system. We find that the ratio of the timescales of the minimum information required by the ML models to accurately capture the dynamics of the SST corresponds to the Reynolds number of the flow. The current framework provides the backbone to explore the capability of such models to capture the dynamics of high-dimensional complex dynamical system like SST flows.

97 MATHEMATICS AND COMPUTING

Understanding latent timescales in neural ordinary differential equation models of advection-dominated dynamical systems

The neural ordinary differential equation (ODE) framework has shown considerable promise in recent years in developing highly accelerated surrogate models for complex physical systems characterized by partial differential equations (PDEs). For PDE-based systems, state-of-the-art neural ODE strategies leverage a two-step procedure to achieve this acceleration: a nonlinear dimensionality reduction step provided by an autoencoder, and a time integration step provided by a neural-network based model for the resultant latent space dynamics (the neural ODE). This work explores the applicability of such autoencoder-based neural ODE strategies for PDEs in which advection terms play a critical role. More specifically, alongside predictive demonstrations, physical insight into the sources of model acceleration (i.e., how the neural ODE achieves its acceleration) is the scope of the current study. Such investigations are performed by quantifying the effects of both autoencoder and neural ODE components on latent system time-scales using eigenvalue analysis of dynamical system Jacobians. To this end, the sensitivity of various critical training parameters – de-coupled versus end-to-end training, latent space dimensionality, and the role of training trajectory length, for example – to both model accuracy and the discovered latent system timescales is quantified. Furthermore, this work specifically uncovers the key role played by the training trajectory length (the number of rollout steps in the loss function during training) on the latent system timescales: larger trajectory lengths correlate with an increase in limiting neural ODE time-scales, and optimal neural ODEs are found to recover the largest time-scales of the full-order (ground-truth) system. Demonstrations are performed across fundamentally different unsteady fluid dynamics configurations influenced by advection: (1) the Kuramoto–Sivashinsky equations (2) Hydrogen-Air channel detonations (the compressible reacting Navier–Stokes equations with detailed chemistry), and (3) 2D Atmospheric flow.

Advection-dominated dynamical systems

Generative learning for slow manifolds and bifurcation diagrams

In dynamical systems characterized by separation of time scales, the approximation of so called “slow manifolds”, on which the long term dynamics lie, is a useful step for model reduction. Initializing on such slow manifolds is a useful step in modeling, since it circumvents fast transients, and is crucial in multiscale algorithms (like the equation-free approach) alternating between fine scale (fast) and coarser scale (slow) simulations. In a similar spirit, when one studies the infinite time dynamics of systems depending on parameters, the system attractors (e.g., its steady states) lie on bifurcation diagrams (curves for one-parameter continuation, and more generally, on manifolds in state parameter space. Sampling these manifolds gives us representative attractors (here, steady states of ODEs or PDEs) at different parameter values. Algorithms for the systematic construction of these manifolds (slow manifolds, bifurcation diagrams) are required parts of the “traditional” numerical nonlinear dynamics toolkit. In more recent years, as the field of Machine Learning develops, conditional score-based generative models (cSGMs) have been demonstrated to exhibit remarkable capabilities in generating plausible data from target distributions that are conditioned on some given label. It is tempting to exploit such generative models to produce samples of data distributions (points on a slow manifold, steady states on a bifurcation surface) conditioned on (consistent with) some quantity of interest (QoI, observable). In this work, we present a framework for using cSGMs to quickly (a) initialize on a low-dimensional (reduced-order) slow manifold of a multi-time-scale system consistent with desired value(s) of a QoI (a “label”) on the manifold, and (b) approximate steady states in a bifurcation diagram consistent with a (new, out-of-sample) parameter value. This conditional sampling can help uncover the geometry of the reduced slow-manifold and/or approximately “fill in” missing segments of steady states in a bifurcation diagram. Finally, the quantity of interest, which determines how the sampling is conditioned, is either known a priori or identified using manifold learning-based dimensionality reduction techniques applied to the training data.

Dynamical systems

Celestial Topology, Symmetry Theories, and Evidence for a NonSUSY D3‐Brane CFT

Symmetry Theories (SymThs) provide a flexible framework for analyzing the global categorical symmetries of a D -dimensional QFT D in terms of a (D + 1)-dimensional bulk system SymTh D+1 . In QFTs realized via local string backgrounds, these SymThs naturally arise from dimensional reduction of the linking boundary geometry. To track possible time dependent effects we introduce a celestial generalization of the standard “boundary at infinity” of a SymTh. As an application of these considerations we revisit large N quiver gauge theories realized by spacetime filling D3-branes probing a non-supersymmetric orbifold $\mathbb{R}$ 6 /Γ. Comparing the imprint of symmetry breaking on the celestial geometry at small and large ‘t Hooft coupling we find evidence for an intermediate symmetry preserving conformal fixed point.

conformal field theory

Bayesian Calibration of Stochastic Agent Based Model via Random Forest

Agent-based models (ABM) provide an excellent framework for modeling outbreaks and interventions in epidemiology by explicitly accounting for diverse individual interactions and environments. However, these models are usually stochastic and highly parametrized, requiring precise calibration for predictive performance. When considering realistic numbers of agents and properly accounting for stochasticity, this high-dimensional calibration can be computationally prohibitive. This paper presents a random forest-based surrogate modeling technique to accelerate the evaluation of ABMs and demonstrates its use to calibrate an epidemiological ABM named CityCOVID via Markov chain Monte Carlo (MCMC). The technique is first outlined in the context of CityCOVID's quantities of interest, namely hospitalizations and deaths, by exploring dimensionality reduction via temporal decomposition with principal component analysis (PCA) and via sensitivity analysis. The calibration problem is then presented, and samples are generated to best match COVID-19 hospitalization and death numbers in Chicago from March to June in 2020. Further, these results are compared with previous approximate Bayesian calibration (IMABC) results, and their predictive performance is analyzed, showing improved performance with a reduction in computation.

60 APPLIED LIFE SCIENCES

Toward memory-efficient melt pool monitoring: a classification framework using event-based imaging and sparse sensing technique

Vision sensors like CMOS and CCD cameras are often used for in-process monitoring of melt pools in laser-based additive and welding processes, but they require transferring large amounts of data and computational processing resources. Event-based neuromorphic imagery, on the other hand, detects only the change in pixel intensity, thus potentially reducing the data amount and latency. With an event imager, this study develops a framework for melt pool condition classification, including image construction, time scale selection, optimal pixel selection, and sparse classification, to achieve a highly memory-efficient scheme. These are based on sparse sensing techniques with singular value decomposition (SVD) and QR pivoting, the two fundamental matrix transformations for linear dimensionality reduction. The framework is then validated by classifying a controlled experiment by exciting various mode shapes of liquid gallium pools of varying depths (3, 6, and 8 mm). At 200 pixels, the classifier can reach overall accuracy of 75%, while at 2000 pixels (0.013% of the total possible pixels), the accuracy is nearly 90% (89.86%). At the same number of pixels, random selection can only achieve 46% and 67%, respectively. The memory savings of the sparsely sampled event data compared to a conventional imager is about 500 times. In addition to performance, implementation and limitations of the framework are also discussed.

42 ENGINEERING

Quantifying Microstructure Variability in Laser Powder Bed Fusion 316 L Stainless Steel Microstructures with Spatial Statistics

Here, we have explored data-driven methods for material microstructure quantification that improve sensitivity to microstructural changes compared to traditional approaches. The methods integrate multiple microstructural properties, including grain morphology, crystallographic orientation, and material phase information. The simpler method employs maps of the Euclidean distance transformation metric to evaluate the morphology of grain boundary networks. The more intensive approach employs generalized spherical harmonic mapping for crystallographic orientations, per-pixel phase information, and a variational auto-encoder for dimensionality reduction and results in a multidimensional clustering of by microstructure similarity. Applied to an experimental dataset of additively manufactured steel, both methods detected slight variations in samples produced under nominally identical processing conditions. Both methods were able to distinguish between samples from multiple (nominally identical) builds, while the generalized spherical harmonics-based method could additionally cluster data samples rotated at two orientations on the build plate. The improved sensitivity of the methods, demonstrated through comparison with traditional microstructure characterization techniques, offers advantages for microstructure quantification and comparisons in advanced manufacturing applications.

SS316L

Projection-based multifidelity linear regression for data-scarce applications

Surrogate modeling for systems with high-dimensional quantities of interest remains challenging, particularly when training data are costly to acquire. This work develops multifidelity methods for multiple-input multiple-output linear regression targeting data-limited applications with high-dimensional outputs. Multifidelity methods integrate many inexpensive low-fidelity model evaluations with limited, costly high-fidelity evaluations. We introduce two projection-based multifidelity linear regression approaches with linear and nonlinear features that leverage principal component basis vectors for dimensionality reduction and combine multifidelity data through: (i) a direct data augmentation using low-fidelity data, and (ii) a data augmentation incorporating explicit linear corrections between low-fidelity and high-fidelity data. The data augmentation approaches combine high-fidelity and low-fidelity data into a unified training set and train the linear regression model through weighted least squares with fidelity-specific weights. We introduce a proximity-based weighting scheme with automatic weight selection strategy through cross-validation. Here, the proposed multifidelity linear regression methods are demonstrated on approximating the surface pressure field of a hypersonic vehicle in flight and the temperature field on an aircraft disc braking system. In an ultra low-data regime of no more than twelve high-fidelity samples, multifidelity linear regression achieves approximately 2% – 12% improvement in median accuracy and a higher R 2 score relative to single-fidelity methods at comparable computational cost.

data augmentation

Karhunen–Loève deep learning method for surrogate modeling and approximate Bayesian parameter estimation

We evaluate the performance of the Karhunen-Loève Deep Neural Network (KL-DNN) framework for surrogate modeling and approximate Bayesian parameter estimation in partial differential equation models. In the surrogate model, the Karhunen-Loève (KL) expansions are used for the dimensionality reduction of the number of unknown parameters and variables, and a deep neural network is employed to relate the reduced space of parameters to that of the state variables. The KL-DNN surrogate model is used to formulate a maximum-a-posteriori-like least-squares problem, which is randomized to draw samples of the posterior distribution of the parameters. We test the proposed framework for a hypothetical unconfined aquifer via comparison with the forward MODFLOW and inverse PEST++ iterative ensemble smoother (IES) solutions as well as the state-of-the-art Fourier neural operator (FNO) and deep operator networks (DeepONets) operator learning surrogate models. Our results show that the KL-DNN surrogate model outperforms FNO and DeepONet for forward predictions. For solving inverse problems, the randomized algorithm provides the same or more accurate Bayesian predictions of the parameters than IES as evidenced by the higher log-predictive probability of both the estimated parameter field and the forecast hydraulic head. The posterior mean obtained from the randomized algorithm is closer to the reference parameter field than that obtained with FNO as the maximum a posteriori estimate.

Approximate Bayesian inference

Personalized and uncertainty-aware coronary hemodynamics simulations: From Bayesian estimation to improved multi-fidelity uncertainty quantification

Non-invasive simulations of coronary hemodynamics have improved clinical risk stratification and treatment outcomes for coronary artery disease, compared to relying on anatomical imaging alone. However, simulations typically use empirical approaches to distribute total coronary flow amongst the arteries in the coronary tree, which ignores patient variability, the presence of disease, and other clinical factors. Further, uncertainty in the clinical data often remains unaccounted for in the modeling pipeline. We present an end-to-end uncertainty-aware pipeline to (1) personalize coronary flow simulations by incorporating vessel-specific coronary flows as well as cardiac function; and (2) predict clinical and biomechanical quantities of interest with improved precision, while accounting for uncertainty in the clinical data. We assimilate patient-specific measurements of myocardial blood flow from clinical CT myocardial perfusion imaging to estimate branch-specific coronary artery flows. Simulated noise in the clinical data is used to estimate the joint posterior distributions of the model parameters using adaptive Markov Chain Monte Carlo sampling. Additionally, the posterior predictive distribution for the relevant quantities of interest is determined using a new approach combining multi-fidelity Monte Carlo estimation with non-linear, data-driven dimensionality reduction. This leads to improved correlations between high- and low-fidelity model outputs. Our framework accurately recapitulates clinically measured cardiac function as well as branch-specific coronary flows under measurement noise uncertainty. We observe substantial reductions in confidence intervals for estimated quantities of interest compared to single-fidelity Monte Carlo estimation and state-of-the-art multi-fidelity Monte Carlo methods. This holds especially true for quantities of interest that showed limited correlation between the low- and high-fidelity model predictions. In addition, the proposed multi-fidelity Monte Carlo estimators are significantly cheaper to compute than traditional estimators, under a specified confidence level or variance. The proposed pipeline for personalized and uncertainty-aware predictions of coronary hemodynamics is based on routine clinical measurements and recently developed techniques for CT myocardial perfusion imaging. The proposed pipeline offers significant improvements in precision and reduction in computational cost.

Bayesian parameter estimation

Online task-space motion control for positioner-coordinated multi-robot manufacturing systems

Incorporating multiple robotic manipulators into large-scale manufacturing systems enhances production efficiency and expands manufacturing capabilities beyond those of single-robot systems. Workpiece positioners in robotic manufacturing have demonstrated significant benefits for process optimization, but coordination strategies for multi-robot systems with shared positioners have received limited attention. This work presents a task-space coordinated trajectory-tracking control framework for multi-robot manufacturing systems, in which robots coordinate their motions within a shared, dynamic workpiece positioning frame. A workpiece positioner actively adjusts the pose of the manufactured component to enable greater operational concurrency and improve overall production efficiency. The proposed motion-coordination scheme employs a distributed and scalable architecture, supporting coordination across heterogeneous multi-robot systems. Two optimization methodologies are introduced to manage kinematic redundancies and maintain continuous, near-optimal operation throughout the manufacturing process. The first strategy exploits a task-space dimensionality reduction to achieve locally optimal configurations by leveraging symmetry-axis rotations of the tool. The second strategy utilizes the workpiece positioner to drive the coordinated robots toward stable and kinematically favorable configurations. For both optimization strategies, multiple objectives are defined to improve key performance metrics, including manipulability, configuration consistency, proximity to mechanical limits, and motion efficiency. Addressing a key limitation of existing coordination approaches, the framework is designed around online setpoint modification, allowing coordinated robots to respond effectively to in-situ process feedback. The proposed control framework is validated using the Robot Operating System (ROS) middleware on a combination of physical and simulated multi-robot system hardware.

Arbogast, Alex [ORNL] (ORCID:0000000154740723)