Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “k mean”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

A dendritic strontium river isoscape for fisheries applications in the Sacramento River basin, California, USA

Objective Understanding the origins and movements of fish is fundamental to effective conservation and fisheries management. Strontium isotope ratios ( 87 Sr/ 86 Sr) in otoliths provide a powerful tracer of natal origin and migratory pathways. However, existing 87 Sr/ 86 Sr isoscapes for the Sacramento River basin, an ecosystem that supports ecologically and economically important salmon populations, rely on discrete classification approaches that overlook unsampled habitats and do not incorporate spatial uncertainty. Our objective was to develop a continuous, network-explicit 87 Sr/ 86 Sr isoscape with quantified uncertainty to fill in data gaps and enable probabilistic assignments of fish origin and movement. Methods We used river water 87 Sr/ 86 Sr data from 106 sites (1997–2021) to develop spatial stream network models that use dendritic connectivity and watershed characteristics (lithology, bedrock age, and land cover) to predict river water 87 Sr/ 86 Sr throughout the basin. Models were fitted using maximum and restricted likelihood and were evaluated via Akaike’s information criterion and leave-one-out cross validation. We produced both historical (pre-dam) and present-day (below-dam) isoscapes, delineated uncertainty-informed isotopic ranges using k -means clustering, and applied a proof-of-concept Bayesian assignment to estimate natal origins and early rearing habitats for two endangered winter-run Chinook Salmon Oncorhynchus tshawytscha. Results Cross validation indicated strong performance of the 87 Sr/ 86 Sr model (leave-one-out cross validation: R 2 = 0.91; root mean square error = 0.0005). Uncertainty-informed clustering identified 19 isotopic “suites” (reaches with indistinguishable 87 Sr/ 86 Sr values) in present-day anadromous habitats and 25 suites in the historical network. Example natal and early rearing assignments included predictions that challenged expectations for juvenile salmon migration based on predicted river 87 Sr/ 86 Sr compositions. Conclusions This study developed a continuous, network-explicit 87 Sr/ 86 Sr isoscape that integrates existing river data to predict 87 Sr/ 86 Sr in unsampled reaches and the likely achievable range and resolution of otolith-based origin and life history inference. The resulting river isoscape provides a valuable tool to predict salmon movements and identify habitats supporting their survival and growth that otherwise might remain undetected. Coupling these predictions with complementary approaches that ground-truth juvenile presence (e.g., targeted fish surveys) represents an important step toward science-informed restoration and management of critical habitats throughout the Sacramento River basin.

Environmental sciences↗

ThickBrick: optimal event selection and categorization in high energy physics. Part I. Signal discovery

We provide a prescription called ThickBrick to train optimal machine-learning-based event selectors and categorizers that maximize the statistical significance of a potential signal excess in high energy physics (HEP) experiments, as quantified by any of six different performance measures. For analyses where the signal search is performed in the distribution of some event variables, our prescription ensures that only the information complementary to those event variables is used in event selection and categorization. This eliminates a major misalignment with the physics goals of the analysis (maximizing the significance of an excess) that exists in the training of typical ML-based event selectors and categorizers. In addition, this decorrelation of event selectors from the relevant event variables prevents the background distribution from becoming peaked in the signal region as a result of event selection, thereby ameliorating the challenges imposed on signal searches by systematic uncertainties. Our event selectors (categorizers) use the output of machine-learning-based classifiers as input and apply optimal selection cutoffs (categorization thresholds) that are functions of the event variables being analyzed, as opposed to flat cutoffs (thresholds). These optimal cutoffs and thresholds are learned iteratively, using a novel approach with connections to Lloyd’s k-means clustering algorithm. We provide a public, Python implementation of our prescription, also called ThickBrick, along with usage examples.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Global teleconnections influencing large-scale drought in the United States using SVDI

Understanding recent large-scale drought patterns and the mechanisms producing extreme drought events is vital for future drought forecasts and understanding future drought risks. Increasingly, vapor pressure deficit (VPD) has been used as an important measure of evaporative demand and proxy for drought detection. In this study, VPD is used to calculate the new Standardized VPD Drought Index (SVDI) with NASA North American Land Data Assimilation System (NLDAS) data. Previous studies have shown that SVDI accurately identifies the timing and magnitude short-term droughts in the United States (U.S). In the present study, SVDI is now used to identify large-scale drought patterns between 1980 and 2021 and drought variability driven by selected global teleconnections originating in the Pacific and Atlantic Oceans. Spatial drought characteristics were extracted from SVDI using empirical orthogonal function (EOF) analysis. Then a k-means clustering algorithm was applied to both EOF principal components and primary teleconnections, including the El Nino-Southern Oscillation (ENSO) and Pacific Decadal Oscillation (PDO) to identify drought events driven by the Pacific Ocean. Results show that the SVDI is useful in evaluating large-scale drought variability in the U.S. related to global teleconnections, and that mechanisms influencing summer drought patterns in the Western and Southwestern U.S. are driven by a tropical-extratropical interactions originating in the equatorial Pacific Ocean related to ENSO dynamics with interdecadal variability modulated by PDO. The large-scale droughts in the Central and Southern U.S., like those in 2011 and 2012, on the other hand, are driven by the North Pacific Ocean warm pool during a strong negative PDO, which subsequently influenced variability in the Bermuda-Azores High in the Atlantic Ocean. In summer 2011, the Bermuda-Azores High weakened, reducing the onshore winds and moisture transport along the eastern Gulf of Mexico and contributing to ongoing drought in the region. The Northern Pacific and Atlantic Ocean sea surface temperatures (SSTs) have increased between 1980 and 2021. In conclusion, as SSTs continue to rise in the Northern Pacific Ocean, one consequence of the coupled North Pacific warm pool and atmospheric dynamics, is to increase summer drought variability over a large region in the southern and midwestern U.S. under global warming.

54 ENVIRONMENTAL SCIENCES↗

Modeling regional precipitation over the Indus River basin of Pakistan using statistical downscaling

Complex processes govern spatiotemporal distribution of precipitation within the high-mountainous headwater regions (commonly known as the upper Indus basin (UIB)), of the Indus River basin of Pakistan. Reliable precipitation simulations particularly over the UIB present a major scientific challenge due to regional complexity and inadequate observational coverage. Here, we present a statistical downscaling approach to model observed precipitation of the entire Indus basin, with a focus on UIB within available data constraints. Taking advantage of recent high altitude (HA) observatories, we perform precipitation regionalization using K-means cluster analysis to demonstrate effectiveness of low-altitude stations to provide useful precipitation inferences over more uncertain and hydrologically important HA of the UIB. We further employ generalized linear models (GLM) with gamma and Tweedie distributions to identify major dynamic and thermodynamic drivers from a reanalysis dataset within a robust cross-validation framework that explain observed spatiotemporal precipitation patterns across the Indus basin. Final statistical models demonstrate higher predictability to resolve precipitation variability over wetter southern Himalayans and different lower Indus regions, by mainly using different dynamic predictors. The modeling framework also shows an adequate performance over more complex and uncertain trans-Himalayans and the northwestern regions of the UIB, particularly during the seasons dominated by the westerly circulations. However, the cryosphere-dominated trans-Himalayan regions, which largely govern the basin hydrology, require relatively complex models that contain dynamic and thermodynamic circulations. Furthermore, we also analyzed relevant atmospheric circulations during precipitation anomalies over the UIB, to evaluate physical consistency of the statistical models, as an additional measure of reliability. Overall, our results suggest that such circulation-based statistical downscaling has the potential to improve our understanding towards distinct features of the regional-scale precipitation across the upper and lower Indus basin. Additionally, such understanding should help to assess the response of this complex, data-scarce, and climate-sensitive river basin amid future climatic changes, to serve communal and scientific interests.

54 ENVIRONMENTAL SCIENCES↗

Root size and soil physicochemical properties drive microscale spatial patterns of Fe and As retention in the rice rhizosphere

Background and Aims: Radial oxygen loss from rice roots in flooded soils oxidizes and precipitates dissolved Fe(II), Mn(II), and As(III) into mixed Fe(III), Mn(III/IV), and As(V) as root plaque and in the rhizosphere soil. It is unknown how different soils and root sizes impact the spatial extent of Fe and As retention outside the root. Methods: We imaged cross-sections of 90 roots from 6 different soils using synchrotron μXRF imaging followed by k-means clustering and elliptical averaging to distinguish bulk soil, rhizosphere, plaque, and roots based on As and Fe patterns. Results: We found preferential As retention in the plaque and rhizospheres of most roots except small (< 0.45 mm) roots in silty soils with low P or high As. In contrast, clayey soils had similar As-Fe correlations across plaque, rhizosphere, and bulk soil. Large (> 0.45 mm) roots often had no oxidized rhizosphere region. We obtained an extensive dataset of 256 As and 155 Mn synchrotron μXANES measurements, which revealed that rhizosphere and plaque As was mainly inorganic As(V) and As(III), and Mn oxidation state varied between soils but not between belowground locations. Conclusion: Small roots in coarse-textured soils were less likely to have As retention in the plaque or rhizosphere compared to large roots and fine-textured soils. Furthermore, the unique and extensive data in this study provides new insight into soil and root size impacts on As retention in the rhizosphere. It is essential to investigate a representative number of samples to draw conclusions from XRF imaging.

36 MATERIALS SCIENCE↗

Predicting the oxidation states of Mn ions in the oxygen-evolving complex of photosystem II using supervised and unsupervised machine learning

Abstract Serial Femtosecond Crystallography at the X-ray Free Electron Laser (XFEL) sources enabled the imaging of the catalytic intermediates of the oxygen evolution reaction of Photosystem II (PSII). However, due to the incoherent transition of the S-states, the resolved structures are a convolution from different catalytic states. Here, we train Decision Tree Classifier and K-means clustering models on Mn compounds obtained from the Cambridge Crystallographic Database to predict the S-state of the X-ray, XFEL, and CryoEM structures by predicting the Mn’s oxidation states in the oxygen-evolving complex. The model agrees mostly with the XFEL structures in the dark S 1 state. However, significant discrepancies are observed for the excited XFEL states (S 2 , S 3, and S 0 ) and the dark states of the X-ray and CryoEM structures. Furthermore, there is a mismatch between the predicted S-states within the two monomers of the same dimer, mainly in the excited states. We validated our model against other metalloenzymes, the valence bond model and the Mn spin densities calculated using density functional theory for two of the mismatched predictions of PSII. The model suggests designing a more optimized sample delivery and illumiation systems are crucial to precisely resolve the geometry of the advanced S-states to overcome the noncoherent S-state transition. In addition, significant radiation damage is observed in X-ray and CryoEM structures, particularly at the dangler Mn center (Mn4). Our model represents a valuable tool for investigating the electronic structure of the catalytic metal cluster of PSII to understand the water splitting mechanism.

Plant Sciences↗

Data Science Techniques, Assumptions, and Challenges in Alloy Clustering and Property Prediction

Data analytics methods have been increasingly applied to understanding materials chemistry, processing due to the manufacturing approach, and uni-axial and cyclic property relationships in the highly complex space of alloy design. There are several benefits to applying data analytics to this space, including the ability to manage non-linearities in the responses of the alloy attributes and the resulting mechanical properties. However, key difficulties in applying and understanding the results of data analytics include the often lack of reported assumptions and data processing steps necessary to improve interpretation and reproducibility in derived results. In this work, the methods used to generate clustering and correlation analyses for experimental 9% Cr ferritic-martensitic steel data were investigated and the resulting implications for mechanical property predictions were assessed. This work uses principal component analysis, partitioning around medoids, t-SNE, and k-means clustering to investigate trends in composition, processing and microstructure information with creep and tensile properties, building on work done previously using a smaller version of the same dataset. The initial assumptions, preprocessing steps and methods are investigated and outlined in order to depict the fine level of detail required to convey the steps taken to process data and produce analytical results. Here, the variations in the resulting analyses are explored due to the influence of new and more varied data.

36 MATERIALS SCIENCE↗

Assessment of Outliers in Alloy Datasets Using Unsupervised Techniques

We report advancements in data analytics techniques have enabled complex, disparate datasets to be leveraged for alloy design. Identifying outliers in a dataset can reduce noise, identify erroneous and/or anomalous records, prevent overfitting, and improve model assessment and optimization. In this work, two alloy datasets (9-12% Cr ferritic martensitic steels, and austenitic stainless steels) have been assessed for outliers using unsupervised techniques and supplemented with domain knowledge. Principal component analysis and k-means clustering were applied to the data, and points were assessed as outliers based on their distance away from other points in the cluster and from other points in the dataset. The outlier characteristics were investigated to determine both cluster-specific and overall trends in the properties of the outlier points. The approach demonstrated here is extensible to other alloy datasets for outlier identification and evaluation to improve the reliability of machine learning and modeling predictions for advanced alloy design.

36 MATERIALS SCIENCE↗

Day-ahead photovoltaic power production forecasting methodology based on machine learning and statistical post-processing

A main challenge towards ensuring large-scale and seamless integration of photovoltaic systems is to improve the accuracy of energy yield forecasts, especially in grid areas of high photovoltaic shares. The scope of this paper is to address this issue by presenting a unified methodology for hourly-averaged day-ahead photovoltaic power forecasts with improved accuracy, based on data-driven machine learning techniques and statistical post-processing. More specifically, the proposed forecasting methodology framework comprised of a data quality stage, data-driven power output machine learning model development (artificial neural networks), weather clustering assessment (K-means clustering), post-processing output optimisation (linear regressive correction method) and the final performance accuracy evaluation. The results showed that the application of linear regression coefficients to the forecasted outputs of the developed day-ahead photovoltaic power production neural network improved the performance accuracy by further correcting solar irradiance forecasting biases. The resulting optimised model provided a mean absolute percentage error of 4.7% when applied to historical system datasets. Finally, the model was validated both, at a hot as well as a cold semi-arid climatic location, and the obtained results demonstrated close agreement by yielding forecasting accuracies of mean absolute percentage error of 4.7% and 6.3%, respectively. Finally, the validation analysis provides evidence that the proposed model exhibits high performance in both forecasting accuracy and stability.

42 ENGINEERING↗

Smart thermostat data-driven U.S. residential occupancy schedules and development of a U.S. residential occupancy schedule simulator

Occupancy schedule is one of the key inputs in Building Energy Modeling (BEM) to reflect the interaction between buildings and occupants. Over the past decades, standardized occupancy schedules, developed mainly by engineering rule-of-thumb, have been widely used in BEM due to its simplicity and lack of real measured occupancy data. However, the BEM community has recognized their association with uncertainty and reliability in simulation results from BEM. This study introduces representative occupancy schedules in the U.S. residential buildings, derived from a large smart thermostat dataset and time-series K-means clustering, and an open-source tool to generate a stochastic residential occupancy schedule. Over 90,000 residential occupancy schedules were estimated from the ecobee Donate Your Data dataset. Then, the representative occupancy schedules were identified through clustering. This study further investigated the impacts of three parameters (day, house type, and state) on residential occupancy schedules. Then, a tool, the Residential Occupancy Schedule Simulator (ROSS), is developed using the representative occupancy schedules derived in this study. Details of this tool are presented in this paper. In conclusion, the derived representative occupancy schedules and the ROSS tool can help improve the energy modeling of residential buildings.

42 ENGINEERING↗

Quantitative microstructural investigation of 3D-printed and cast cement pastes using micro-computed tomography and image analysis

Highlights: • Processing alters the microstructure in 3D-printed lamellar cubic element. • A unique patterned and connected pore microstructure network exists. • Micro-channels are connected through micro-pores present at interfaces regions. • Three distinct pore volume domains exists in hardened microstructure. • Volume and frequency of pore, hydrated, and unhydrated clusters are dissimilar. Microstructural phases and mechanical properties of lamellar 3D-printed and cast hardened cement paste (hcp) elements were investigated using a lab-based X-ray microscope at two levels of magnification (0.4× and 4×). K-means clustering was used for quantitative image analysis. The entire volume of intact 3-days-old 3D-printed and cast hcp elements was characterized at 0.4× magnification. Three microstructural features (macro-pores, micro-channels, and interfacial micro-pores) were found to reside in three distinct pore size domains. The largest pores of the 3D-printed element were larger than the largest pores of the reference cast hcp element. Moreover, the smallest pore sizes of the 3D-printed element were found to be smaller than those present in the cast counterparts. Micro-channels were found to be connected to one another through the micro-pores present at interfacial regions, indicating the presence of a uniquely patterned and interconnected pore network. The role of locally weak and porous interfaces on mechanical response and fracture properties is discussed.

36 MATERIALS SCIENCE↗

Generating realistic building electrical load profiles through the Generative Adversarial Network (GAN)

Building electrical load profiles can improve understanding of building energy efficiency, demand flexibility, and building-grid interactions. Current approaches to generating load profiles are time-consuming and not capable of reflecting the dynamic and stochastic behaviors of real buildings; some approaches also trigger data privacy concerns. In this study, we proposed a novel approach for generating realistic electrical load profiles of buildings through the Generative Adversarial Network (GAN), a machine learning technique that is capable of revealing an unknown probability distribution purely from data. The proposed approach has three main steps: (1) normalizing the daily 24-hour load profiles, (2) clustering the daily load profiles with the k-means algorithm, and (3) using GAN to generate daily load profiles for each cluster. The approach was tested with an open-source database – the Building Data Genome Project. We validated the proposed method by comparing the mean, standard deviation, and distribution of key parameters of the generated load profiles with those of the real ones. The KL divergence of the generated and real load profiles are within 0.3 for majority of parameters and clusters. Additionally, results showed the load profiles generated by GAN can capture not only the general trend but also the random variations of the actual electrical loads in buildings. We report the proposed GAN approach can be used to generate building electrical load profiles, verify other load profile generation models, detect changes to load profiles, and more importantly, anonymize smart meter data for sharing, to support research and applications of grid-interactive efficient buildings.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Data-driven occupant-behavior analytics for residential buildings

Many advances have been made in building technology to help save energy, but influencing the behavior of the occupants is still necessary to achieve low-energy use targets. One of the most practical ways to influence and change occupant behaviors is through incentives. Developing incentives for energy-saving and quantifying the impact of occupant behaviors are both active areas of research. Here, we propose a data analytics framework for detecting changes in occupant behaviors, which will help build an analytics feedback loop from behavior impact to incentive design. The framework has two major parts. The first forecasts energy consumption for each occupant, while the second determines a probability distribution for changes in energy consumption. The parts are interchangeable with other existing machine learning and statistical methods. A specific instantiation of the framework, using kernel ridge-regression for forecasting and k-means to find an empirical behavior distribution, is described in detail. An HVAC use-case with 5 different incentivized behaviors is used as an example to show that the framework can detect behavior changes induced by incentives. Furthermore, we show that some simpler behavior-change detection methods do not work, further justifying the use of advanced analytics.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Simulating water dynamics related to pedogenesis across space and time: Implications for four-dimensional digital soil mapping

Digital soil mapping (DSM) relies on machine-learning and geostatistics to represent soil property observations across space. DSM techniques are powerful but often empirical, being limited to the quality and density of point samples. Water dynamics are closely related to soil variability, and the physics that govern water movement are well known. Hydrological properties can hence be simulated by physical models through space and time, unveiling key characteristics about soils. We propose the use of hydrologic models to map soils across the surface (2D), depth (1D), and time (1D)–which provides a 4D approach to digital soil mapping (4DSM). The Distributed Hydrology Soil Vegetation Model (DHSVM) was applied to a watershed currently under pasture. Moisture sensors and wells were installed at different depths in the watershed on summit, sideslope and toeslope positions to validate the model. DHSVM simulations of soil moisture distribution and depth to saturation were performed during the hydrological year (October 2008-September 2009). Clusters of similar pixels based on soil moisture values were determined using Dynamic Time Warping (DTW) to align temporal data and K-means. Clustering was performed both seasonally and for the entire year. Temporal patterns simulated by DHSVM matched measurements given by moisture sensors and wells. Seasonal clusters differed from the annual cluster. Distinct clusters were observed for each season and with depth, showing that spatiotemporal soil variability is lost when statically assessing soils. Spatiotemporal clusters corroborated field observations of fragipan occurrence not explicitly spatially mapped by Soil Survey Geographic Database (SSURGO). If a connection can be made between water and soils, static and dynamic soil variability can be predicted using physically based hydrologic models. Hydrologic models can benefit soil mapping by enabling reliable 4D simulation of water dynamics, which are fundamental to soil variability and soil classification and directly relate to biological, physical and chemical soil processes not captured by typical soil sampling protocols.

54 ENVIRONMENTAL SCIENCES↗

Exploratory analysis of machine learning techniques in the Nevada geothermal play fairway analysis

Play fairway analysis (PFA) is commonly used to generate geothermal potential maps and guide exploration studies, with a particular focus on locating and characterizing blind geothermal systems. This study evaluates the application of machine learning techniques to PFA in the Great Basin region of Nevada. Following the evaluation of various techniques, we identified two approaches to PFA that produced promising results, 1) supervised Bayesian probabilistic neural networks to generate geothermal potential maps with confidence intervals, and 2) unsupervised principal component analysis paired with k-means clustering to generate both cluster maps to help identify spatial patterns, as well as new combined feature inputs. We applied these techniques to perform a comparative analysis between two principal sets of geological and geophysical features related to permeability and heat and a set of positive (known geothermal resources) and negative training sites (known drill sites with unsuitable geothermal conditions). We found that these methods constrain previously unrecognized feature controls on geothermal favorability, many of which are spatially organized within the extent of cluster groups and the major structural-hydrologic domains of the study area. Furthermore, we utilized exploratory unsupervised modeling to highlight spatial relationships between input data and predictive output results of our supervised modeling. As a result, we demonstrate how our models compare to the previous Nevada PFA and how the rapid insights these machine learning techniques offer may support future assessments of both known and undiscovered blind geothermal systems in the Great Basin region of Nevada and beyond.

15 GEOTHERMAL ENERGY↗

Web-based wide-area monitoring platform for ringdown and clustering analytics in power systems

This paper introduces an open-source research platform for monitoring the Mexican interconnected power grid, allowing real-time processing and information extraction of the grid’s dynamic condition. Moreover, the platform is a Python-based development that embeds different ringdown and clustering analytics tools. In the case of ringdown analysis, the modal information can be extracted using some of the most known algorithms, i.e., Prony analysis, eigensystem realization algorithm (ERA), and matrix pencil (MP). For clustering analysis, the coherent behaviour of generator and non-generator buses is provided by applying recent state-of-the-art techniques such as affinity propagation, K-means, hierarchical agglomerative clustering, and typicality data analysis. The results of up to 93 PMUs show that this open-source platform suits researchers’ and engineers’ power system dynamic analysis requirements.

Clustering↗

A generalizable machine learning-assisted fast Fourier transform algorithm to simulate the large strain phenomena in polycrystalline materials

Machine learning methods have shown initial promise in constitutive modeling for single crystals or homogenized polycrystals, delivering notable computational efficiency. However, existing machine learning-based constitutive models often lack generalizability, limiting their application across diverse boundary value problems. This study introduces a thermodynamics-informed artificial neural network model to accelerate rate-tangent crystal plasticity fast Fourier transform simulations for cross-scale deformation behaviors of polycrystals under complex loading. Our model integrates microstructural variability and local interactions effectively. To address local effects in each grain, we employ K-means clustering to group Gauss points within the microstructure into clusters assumed to be in similar mechanical states. This approach, based on self-clustering analysis, extends model scope from macroscopic stress response to the granular level, capturing mechanical responses and orientation evolution across grains. This reduces the number of nonlinear problems to solve, with cluster responses propagated throughout each group. The thermodynamics-based artificial neural network-extracted features are further processed using local material state clusters to account for history-dependent deformation and evolving microstructures. Additionally, representative volume element simulations with rate-tangent crystal plasticity fast Fourier transform provide reliable datasets for model training. The proposed model demonstrates high efficiency, accuracy, self-consistency, and enhanced generalizability in predicting strain–stress responses and orientation evolution at both individual grain and aggregate scales under complex loading conditions, such as biaxial tension and arbitrary loading scenarios.

36 MATERIALS SCIENCE↗

Polynomial chaos expansions on principal geodesic Grassmannian submanifolds for surrogate modeling and uncertainty quantification

In this work we introduce a manifold learning-based surrogate modeling framework for uncertainty quantification in high-dimensional stochastic systems. Our first goal is to perform data mining on the available simulation data to identify a set of low-dimensional (latent) descriptors that efficiently parameterize the response of the high-dimensional computational model. To this end, we employ Principal Geodesic Analysis on the Grassmann manifold of the response to identify a set of disjoint principal geodesic submanifolds, of possibly different dimension, that captures the variation in the data. Since operations on the Grassmann require the data to be concentrated, we propose an adaptive algorithm based on Riemannian K-means and the minimization of the sample Fréchet variance on the Grassmann manifold to identify “local” principal geodesic submanifolds that represent different system behavior across the parameter space. Polynomial chaos expansion is then used to construct a mapping between the random input parameters and the projection of the response on these local principal geodesic submanifolds. Here, the method is demonstrated on four test cases, a toy-example that involves points on a hypersphere, a Lotka-Volterra dynamical system, a continuous-flow stirred-tank chemical reactor system, and a two-dimensional Rayleigh-Bénard convection problem.

42 ENGINEERING↗