Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Mixture modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Predicting Li-Ion Battery Capacity Fade Using Early-Life Data and a Hybrid Data-Driven Gaussian Process-Bayesian Regression Approach

Accurately predicting Li-ion battery capacity trajectories using early-life data can dramatically improve battery-life understandings and be used to rapidly evaluate design/cost/performance trade-offs when developing new battery materials. Accurate early-life predictions enable researchers to quickly iterate over cell designs and material precursor properties without consistently cycling cells to failure. To this end, we present a toolbox that uses a combined Gaussian Process and Bayesian regression approach that capitalizes on signals other than just capacity (e.g., dQ/dV, voltage drops) to rapidly predict capacity-fade trajectories. The prediction tool uses Bayesian regression to fit functional forms, e.g., power law, sigmoids, etc., to predict capacity-fade dynamics. By fitting functional forms, the capacity fade can be interrogated at any point in the future, allowing for early cell-failure prediction. Additionally, Bayesian regression allows for accurate uncertainty estimates that account for cell-to-cell variability (aleatoric uncertainty) and the lack of observation data (epistemic uncertainty). By only using early cycle data to predict the capacity fade trajectory, uncertainty bounds at end-of-life can be extremely large. The large uncertainty bounds are further exacerbated because there is no systematic way to define the prior distribution of the functional forms' parameters. We improve our the predicted trajectory confidence interval of our predicted trajectory using two methods. First, we shows that a small amount of held-out cycling data is sufficientuse some train cells, that have been cycled to failure to derive information regarding the appropriate prior distributions for the functional forms' parameters of the functional form, effectively leading to data-driven priors.. We propose constructing the data-driven priors by first running a Bayesian regression starting with uninformed priors to generate intermediate cell-specific posterior parameter distributions. These posterior distributions are combined using a Ggaussian mixture model for each parameter to create the data-driven priors. These mixture models serve as the data-driven prior distributions for the parameters for. Second, we derive multiple features, e.g., C_dchg 0.5 DoD 0.5, log (|mean(dQ/dV_(w_3-w_0 ) (V)|), etc., from the train cellsheld-out cycling data, identify which the features are that best predicting capacity at early/mid-life cycles, and then create Ggaussian process regression models that are used for predicting capacity at early/mid-life cycles for the test cells (see blue dots with error bars in Fig 1b). Finally, these predicted data-points are used in addition to the actual early cycle data capacity fade to construct the Bayesian regression trajectory for the test cell s. Notably. We note that these two methods are complementary and can be combined with each other. We evaluate the performance of our proposed method on an testing open-source dataset from Iowa State University and Iowa Lakes Community College (ISU-ILCC). This dataset comprises of 251 nickel-manganese-cobalt/graphite Lithium-ion cells that are cycled under 63 different conditions. We compute the mean average percentage error (MAPE) and negative log predictive density (NLPD) to quantify the efficacy of our method. Our initial findings suggest that, when only few observations are available, for test cells, when using only Bayesian regression with uninformed priors, a power law functional provides the most accurate predictions. with very few data points. However, asHowever, a the number of data points increases, a twin sigmoidal function becomes more accurate as the number of observations further increases. We also find that using as little as 10% of the data set towards generating data-driven priors can lead to significant improvement in prediction accuracy when using early cycle data. Lastly, we found that augmenting early-cycle data with Gaussian process-predicted capacity data for Bayesian regression greatly improves the prediction accuracy. We will present a comprehensive comparison of our methods to other methods available in the literature and apply this method to additional battery datasets.

42 ENGINEERING↗

Multi-dimensional modeling of mixture preparation in a direct injection engine fueled with gaseous hydrogen

With the recent advances of direct injection (DI) technology, introducing hydrogen into the combustion chamber through DI is being considered as a viable approach to circumvent backfire and pre-ignition encountered in early generations of hydrogen engines. As part of a broader vision to develop a robust numerical model to study hydrogen spark ignition (SI) combustion in internal combustion (IC) engines, the present numerical investigation focuses on mixture preparation in a hydrogen DI SI engine. This study is carried out with a single hole injector with gaseous hydrogen injected at 100 bar injection pressure. Simulations are carried out for high and low tumble configurations and validated against optical data acquired from planar laser induced fluorescence (PLIF) measurements. Varying mesh configurations are investigated for the impact on in-cylinder mixture distribution. A particular emphasis is placed on the effect of nozzle geometry and mesh orientation near the wall. Overall, the computational model is found to predict the mixture distribution in the combustion cylinder reasonably well. The results showed that the alignment of mesh with the flow direction is important to achieve good agreement between numerical analysis and optical measurement data.

08 HYDROGEN↗

Observed and Imputed Volumetric Soil Water Content Timeseries for the New Mexico Elevation Gradient

Reliable soil water content (SWC) data are essential for understanding dryland ecosystem dynamics, but high-frequency SWC sensors often fail, creating gaps in critical datasets. To address this, we developed a Bayesian mixture model that imputes missing SWC using both linear interpolation and an ecosystem water balance model (SOILWAT2), tested across six AmeriFlux eddy covariance tower sites in the New Mexico Elevation Gradient, demonstrating its effectiveness in reconstructing SWC patterns while providing insights into the factors driving SWC variability. Daily volumetric soil water content (SWC) data are provided as csv-formatted spreadsheets for the six AmeriFlux sites (US-Seg, US-Ses, US-Wjs, US-Mpi, US-Vcp, and US-Vcs). For each site there is an observed SWC file (site_SWC_gapfill.csv) and a file that contains imputed SWC (imputed_SWC_site.csv). The observed SWC files contain temperature corrected sensor values, tower precipitation data, as well as outputs from SOILWAT2 simulations that were used to impute SWC. The imputed files contain the original observed SWC values and the imputed missing SWC values. When SWC was missing from the original data, the missing value was imputed based on the Bayesian imputation mixture model. The posterior mean of all imputed values is reported as "mean_X". When the observed SWC was NOT missing, mean_X = observed SWC value (original data). The standard deviation, 2.5th percentile and the 97.5th percentile for the imputed values are also reported in the imputed files. There are readme text files for each file type explaining the contents of each column.

54 ENVIRONMENTAL SCIENCES↗

A unified exploration of the chronology of the Galaxy

The Milky Way has distinct structural stellar components linked to its formation and subsequent evolution, but disentangling them is non-trivial. With the recent availability of high-quality data for a large numbers of stars in the Milky Way, it is a natural next step for research in the evolution of the Galaxy to perform automated explorations with unsupervised methods of the structures hidden in the combination of large-scale spectroscopic, astrometric, and asteroseismic data sets. We determine precise stellar properties for 21 076 red giants, mainly spanning 2–15 kpc in Galactocentric radii, making it the largest sample of red giants with measured asteroseismic ages available to date. We explore the nature of different stellar structures in the Galactic disc by using Gaussian mixture models as an unsupervised clustering method to find substructure in the combined chemical, kinematic, and age subspace. The best-fitting mixture model yields four distinct physical Galactic components in the stellar disc: the thin disc, the kinematically heated thin disc, the thick disc, and the stellar halo. We find hints of an age asymmetry between the Northern and Southern hemisphere, and we measure the vertical and radial age gradient of the Galactic disc using the asteroseismic ages extended to further distances than previous studies.

79 ASTRONOMY AND ASTROPHYSICS↗

Development and validation of the cavitation-induced erosion risk assessment tool

This work presents the development of a cavitation-induced erosion risk assessment (CIERA) tool that links multiphase flow simulation predictions with the progress towards material erosion. To develop a robust erosion modeling tool, the cavitation and erosion predictions for pressurized diesel fuel flow within a channel geometry were validated over a range of Reynolds and cavitation number conditions in two different aluminum channel geometries, one featuring a rounded inlet corner and the other with a sharp inlet corner. The multiphase flow development within the channel was modeled using a compressible mixture model, where phase change was represented with the homogeneous relaxation model and the turbulent flow evolution was modeled using a dynamic structure approach for Large Eddy Simulations. To improve representation of the incubation period before material rupture over existing approaches, a physics-based metric was derived based on the cumulative energy absorbed by the solid material from repeated hydrodynamic impacts. When the average peak pressure was related to the incubation period, the incubation period and its sensitivity to changes in flow conditions were found to be overpredicted. In contrast, predictions from CIERA provided a more accurate means to qualitatively and quantitatively predict the influence of flow conditions on the incubation period before material erosion. When the predicted stored energy was related to the solid material properties to estimate the incubation period, multiphase flow simulations demonstrated accurate representation of the sensitivity of erosion severity to changes in flow conditions. The use of CIERA led to quantitative agreement of the predicted incubation period within 5% of the experimentally measured incubation period.

42 ENGINEERING↗

Thermal Overloading Risk Mitigation With a Semi-Analytical Probabilistic Model on Branch Current

A semi-analytical formulation is presented in this paper for the probability computation of branch current in multiphase systems. The developed formula is derived based on the linear power flow model in rectangular coordinates. The system uncertainty injections can be renewable energy resources or loads and are modeled using a Gaussian mixture model (GMM). The developed formula can be used to compute the line current violation probability as well as integrate into optimal power flow problem as chance-constraint relaxation. Here, the proposed formula is first compared with the Matlab embedded numerical integration function to show its performance. Besides, the semi-analytical formula is validated and compared with the Monte Carlo simulation method using the IEEE 123-bus system, EPRI Ckt5, and Ckt7 systems.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Unsupervised probabilistic models for sequential Electronic Health Records

We develop an unsupervised probabilistic model for heterogeneous Electronic Health Record (EHR) data. Utilizing a mixture model formulation, our approach directly models sequences of arbitrary length, such as medications and laboratory results. This allows for subgrouping and incorporation of the dynamics underlying heterogeneous data types. The model consists of a layered set of latent variables that encode underlying structure in the data. These variables represent subject subgroups at the top layer, and unobserved states for sequences in the second layer. We train this model on episodic data from subjects receiving medical care in the Kaiser Permanente Northern California integrated healthcare delivery system. The resulting properties of the trained model generate novel insight from these complex and multifaceted data. In addition, we show how the model can be used to analyze sequences that contribute to assessment of mortality likelihood.

59 BASIC BIOLOGICAL SCIENCES↗

Systems Engineering of Rhodococcus opacus to Enable Production of Drop-in Fuels from Lignocellulose

Production of drop-in fuels from lignocellulose using Rhodococcus opacus PD630 (hereafter R. opacus) is a challenging goal. During the grant period we have pushed the field forward significantly in several areas of research. Towards the end goal of accelerating the adoption of R. opacus in biofuel production, during the grant period we have expanded the phenotypic characterization of R. opacus grown in single aromatic (model lignocellulosic) compounds or their mixtures, modeling the growth conditions in lignocellulosic biomass. Harnessing the power of adaptive evolution, we produced evolved R. opacus isolates with superior lignin valorization capabilities and identified differentially expressed genes and pathways after adaptation. We used next generation multi-omic techniques such as genomic, transcriptomic, and metabolomic analyses, to identify the catabolic pathways used by R. opacus to degrade aromatic compounds and funnel these degradation products into central metabolism, as well as the aromatic transport genes required for increased tolerance and utilization. Taking this information one step further, we identified endogenous transcription factors and regulatory mechanisms important for degradation of five model aromatic compounds. To accurately estimate R. opacus growth and consumption on model lignin compounds we pioneered the use of novel extraction procedures prior to GC-MS analysis. Alongside 13 C-metabolic flux analysis, we have elucidated the metabolic routes preferred by Rhodococcus opacus during aromatic compound degradation. Finally, we used in tandem lipidomics and high-resolution mass spectrometry to identify the modulation of mycolic acids and phospholipid membrane composition modification as a strategy for aromatic tolerance in R. opacus. Being a non-model organism, R. opacus lacks the breadth of tools and technical foundation which drive biofuel research in more well-understood microbes such as Escherichia coli. To reduce this burden for use, we designed and produced new tools for genomic manipulation and engineering in R. opacus. These engineering breakthroughs support efficient genomic editing, enabling gene overexpression, repression, and genetic alteration. Using these tools, we have generated synthetically engineered strains with increased lipogenesis and growth, both positive traits required for increased lignin valorization. Optimizing engineered strains for biofuel production from lignocellulose requires extremely sophisticated synthetic rewiring of metabolism. To facilitate systems-level reorganization of metabolism in R. opacus, we created a genome-scale model that accurately predicts metabolic flux and growth rates on the aromatic compound phenol. Lignin requires extensive pre-treatment before biological degradation by R. opacus. Towards an eventual goal of degrading real-world lignin, we developed new depolymerization processes to generate lignin breakdown products (LBP). We optimized LBP storage and composition analysis techniques, enabling accurate prediction of specific LBP compound integration into cell wall components. Overall, through the work funded by this grant we generated 20 manuscripts (17 published, 3 in review/preparation), methods for increased accuracy in metabolomics of aromatic compounds, multiple genetic tools for altering the R. opacus genome, genome scale models for predicting flux through metabolic pathways, as well as multi-omic data for community use. The work funded by this grant has increased the knowledge of aromatic degradation in bacteria and advanced our efforts to optimize R. opacus for lignin valorization.

09 BIOMASS FUELS↗

Nonlinear sparse Bayesian learning for physics-based models

This paper addresses the issue of overfitting while calibrating unknown parameters of over-parameterized physics-based models with noisy and incomplete observations. Here, a semi-analytical Bayesian framework of nonlinear sparse Bayesian learning (NSBL) is proposed to identify sparsity among model parameters during Bayesian inversion. NSBL offers significant advantages over machine learning algorithm of sparse Bayesian learning (SBL) for physics-based models, such as 1) the likelihood function or the posterior parameter distribution is not required to be Gaussian, and 2) prior parameter knowledge is incorporated into sparse learning (i.e. not all parameters are treated as questionable). NSBL employs the concept of automatic relevance determination (ARD) to facilitate sparsity among questionable parameters through parameterized prior distributions. The analytical tractability of NSBL is enabled by employing Gaussian ARD priors and by building a Gaussian mixture-model approximation of the posterior parameter distribution that excludes the contribution of ARD priors. Subsequently, type-II maximum likelihood is executed using Newton's method whereby the evidence and its gradient and Hessian information are computed in a semi-analytical fashion. We show numerically and analytically that SBL is a special case of NSBL for linear regression models. Subsequently, a linear regression example involving multimodality in both parameter posterior pdf and model evidence is considered to demonstrate the performance of NSBL in cases where SBL is inapplicable. Next, NSBL is applied to identify sparsity among the damping coefficients of a mass-spring-damper model of a shear building frame. These numerical studies demonstrate the robustness and efficiency of NSBL in alleviating overfitting during Bayesian inversion of nonlinear physics-based models.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Spatial patterns in occupancy and density of larval lampreys in freshwater habitats restored to a Stage 0 condition

Abstract We examined occupancy and density of larval lampreys ( Entosphenus tridentatus and Lampetra spp.) in two rivers in Oregon (USA) restored to a Stage 0 condition 1–5 years prior, using a multiscale occupancy model and a zero‐inflated Poisson mixture model. We sampled lampreys using backpack electrofishing in randomly distributed, paired, 1‐m 2 quadrats and recorded environmental data. Probabilities of occupancy and density were higher when water velocity was low, the substrate was noncompacted, and sediment was dominated by fines (<4 mm). At mean water depth (0.34 m) and velocity (0.09 m/s), estimated densities in occupied quadrats were 4.8 lampreys/m 2 (95%: 3.4–6.9) when the substrate was compacted, and fines were not dominant, and 21.1 lampreys/m 2 (95%: 17.7–25.3) when the substrate was noncompacted and fines were dominant. Probabilities of detecting occupancy in a 1‐m 2 quadrat sampled by backpack electrofishing were 0.76 (95%: 0.64–0.87) when captured after visual observation and 0.80 (95%: 0.71–0.88) with blind sweeps (i.e., constantly moving the net regardless of observation). The probability of capturing a single lamprey in a quadrat sampled by blind sweeps was 0.32 (95%: 0.27–0.37). Sampling in paired 1‐m 2 quadrats facilitated concurrent examination of patterns in occupancy and density while accounting for capture probability, which could aid temporal monitoring of restored habitats. To the best of our knowledge, this is the first study to document occupancy and estimate densities of larval lampreys in habitats that underwent valley floor restoration to Stage 0. We observed both lamprey genera within 5 years of restoration. Aquatic restoration that increases low‐velocity, noncompacted, fine sediment habitats could benefit lampreys.

Harris, Julianne E.↗

Model and remote-sensing-guided experimental design and hypothesis generation for monitoring snow-soil–plant interactions

In this study, we develop a machine-learning (ML)-enabled strategy for selecting hillslope-scale ecohydrological monitoring sites within snow-dominated mountainous watersheds, with a particular focus on snow-soil–plant interactions. Data layers rely on spatial data layers from both remote sensing and hydrological model simulations. Specifically, a Landsat-based foresummer drought sensitivity index is used to define the dependency of the annual peak plant productivity on the Palmer drought severity index in the early growing season. Hydrological simulations provide the spatiotemporal dynamics of near-surface soil moisture and snow depth. In this framework, a regression analysis identifies the key hydrological variables relevant to the spatial heterogeneity of drought sensitivity. We then apply unsupervised clustering to these key variables, using the Gaussian mixture model, to group hillslopes into several zones that have divergent relationships regarding soil moisture, snow dynamics, and drought sensitivity. Using the datasets collected in the East River Watershed (Crested Butte, Colorado, United States), results show that drought sensitivity is significantly correlated with model-derived soil moisture and snow-free timing over space and time. The relationship is, however, non-linear, such that the correlation decreases above a threshold elevation and in a heavy snow year due to large snowpacks, lateral flow, and soil storage limitations. Clustering is then able to define the zones that have high or low sensitivity to drought, as well as the mid-elevation regions where sensitivity is associated with the topographic aspect and net potential radiation. In addition, the algorithm identifies the most representative hillslopes with road/trail access within each zone for installing monitoring sites. Our method also aims to significantly increase the use of ML and model-simulation results to guide critical zone and watershed monitoring activities.

54 ENVIRONMENTAL SCIENCES↗

Statistical Behavior of Low-Amplitude Power System Point-on-Wave Measurements

The power grid is undergoing massive changes to ensure resiliency and reliability in a more decentralized world. Distributed energy resources are becoming a prominent source of generation, potentially leading to a lack of centralized generation sources. Due to these new behaviors and system topologies, it is important to install measurement devices that are 1) accurate and 2) self-aware of their measurement quality. In this paper, a residential-scale microgrid is used to generate voltage and current waveforms, captured by Verivolt and National Instruments measurement equipment. A least-squares approach is used to separate the “clean” signals from the noise. Finally, Gaussian mixture modeling is used to approximate noise distributions, and it is shown these higher-order distribution estimates are a better fit to voltage and current noise profiles than single-mode Gaussian estimates.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A Swing of the Pendulum: The Chemodynamics of the Local Stellar Halo Indicate Contributions from Several Radial Merger Events

We find that the chemical abundances and dynamics of APOGEE and GALAH stars in the local stellar halo are inconsistent with a scenario in which the inner halo is primarily composed of debris from a single massive, ancient merger event, as has been proposed to explain the Gaia-Enceladus/Gaia Sausage (GSE) structure. The data contain trends of chemical composition with energy that are opposite to expectations for a single massive, ancient merger event, and multiple chemical evolution paths with distinct dynamics are present. We use a Bayesian Gaussian mixture model regression algorithm to characterize the local stellar halo, and find that the data are fit best by a model with four components. We interpret these components as the Virgo Radial Merger (VRM), Cronus, Nereus, and Thamnos; however, Nereus and Thamnos likely represent more than one accretion event because the chemical abundance distributions of their member stars contain many peaks. Although the Cronus and Thamnos components have different dynamics, their chemical abundances suggest they may be related. We show that the distinct low- and high-α halo populations from Nissen & Schuster are explained by VRM and Cronus stars, as well as some in situ stars. Because the local stellar halo contains multiple substructures, different popular methods of selecting GSE stars will actually select different mixtures of these substructures, which may change the apparent chemodynamic properties of the selected stars. We also find that the Splash stars in the Solar region are shifted to higher v $\phi$ and slightly lower [Fe/H] than previously reported.

79 ASTRONOMY AND ASTROPHYSICS↗

Disaggregating Customer-Level Behind-the-Meter PV Generation Using Smart Meter Data and Solar Exemplars

Customer-level rooftop photovoltaic (PV) has been widely integrated into distribution systems. In most cases, PVs are installed behind-the-meter (BTM), and only the net demand is recorded. Therefore, the native demand and PV generation are unknown to utilities. Separating native demand and solar generation from net demand is critical for improving grid-edge observability. In this paper, a novel approach is proposed for disaggregating customer-level BTM PV generation using low-resolution but widely available hourly smart meter data. The proposed approach exploits the strong correlation between monthly nocturnal and diurnal native demands and the high similarity among PV generation profiles. First, a joint probability density function (PDF) of monthly nocturnal and diurnal native demands is constructed for customers without PVs, using Gaussian mixture modeling (GMM). Deviation from the constructed PDF is utilized to probabilistically assess the monthly solar generation of customers with PVs. Then, to identify hourly BTM solar generation for these customers, their estimated monthly solar generation is decomposed into an hourly timescale; to do this, we have proposed a maximum likelihood estimation (MLE)-based technique that utilizes hourly typical solar exemplars. Leveraging the strong monthly native demand correlation and high PV generation similarity enhances our approach's robustness against the volatility of customers’ hourly load and enables highly-accurate disaggregation. Furthermore, the proposed approach has been verified using real native demand and PV generation data.

14 SOLAR ENERGY↗

G-Mapper: Learning a Cover in the Mapper Construction

The Mapper algorithm is a visualization technique in topological data analysis (TDA) that outputs a graph reflecting the structure of a given dataset. However, the Mapper algorithm requires tuning several parameters in order to generate a “nice” Mapper graph. This paper focuses on selecting the cover parameter. We present an algorithm that optimizes the cover of a Mapper graph by splitting a cover repeatedly according to a statistical test for normality. Our algorithm is based on G-means clustering, which searches for the optimal number of clusters in 𝑘-means by iteratively applying the Anderson–Darling test. Our splitting procedure employs a Gaussian mixture model to carefully choose the cover according to the distribution of the given data. In conclusion, experiments for synthetic and real-world datasets demonstrate that our algorithm generates covers so that the Mapper graphs retain the essence of the datasets, while also running significantly faster than a previous iterative method.

G-means clustering↗

Estimating Higher-Order Moments Using Symmetric Tensor Decomposition

In this paper, we consider the problem of decomposing higher-order moment tensors, i.e., the sum of symmetric outer products of data vectors. Such a decomposition can be used to estimate the means in a Gaussian mixture model and for other applications in machine learning. The dth-order empirical moment tensor of a set of p observations of n variables is a symmetric d-way tensor. Our goal is to nd a low-rank tensor approximation comprising r $\ll$ p symmetric outer products. The challenge is that forming the empirical moment tensor costs O(pn d ) operations and O(n d ) storage, which may be prohibitively expensive; additionally, the algorithm to compute the low-rank approximation costs O(n d ) per iteration. Our contribution is avoiding formation of the moment tensor, computing the low-rank tensor approximation of the moment tensor implicitly using O(pnr) operations per iteration and no extra memory. This advance opens the door to more applications of higher-order moments since they can now be efficiently computed. We present numerical evidence of the computational savings and show an example of estimating the means for higher-order moments.

97 MATHEMATICS AND COMPUTING↗