Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “k mean”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Forming a database to study reversed magnetic shear from the National Spherical Torus eXperiment using machine learning

Achieving a long-lived reversed magnetic shear (RMS) target plasma in the National Spherical Torus eXperiment Upgrade will require developing various sustainment scenarios. To help with the ongoing plasma control efforts, the development of a new analysis for the motional Stark effect (MSE) diagnostic using a machine learning algorithm, namely, MSE-ML, is described. MSE-ML will be used to identify patterns during RMS discharges, some of which suffer magnetohydrodynamic (MHD) events resulting in current redistribution and monotonic q-profiles. A database consisting of q and magnetic shear profiles is being constructed primarily based on the existing National Spherical Torus eXperiment data with equilibrium reconstructions constrained by the magnetic field pitch angle profile measured using the multi-channel MSE diagnostic. An unsupervised k-means clustering of the data is developed to study the RMS formation as a function of time. The initial clustering from the q-profiles shows significant differences in both amplitude and the duration of the RMS period. As a goal, the clustering results that detect and distinguish shots with substantial and sustained RMS are to be used as a preprocessing step in a supervised algorithm to identify the underlying conditions that lead to long-lasting improved confinement with RMS. Another aim of the MSE-ML study is to identify precursors of RMS-destroying MHD events in either derived data such as the q-profile or directly measured data such as the magnetic field pitch angle profile.

Uzun-Kaymak, I. U. (ORCID:0000000276251493)↗

Deterministic High-Fidelity Neutronics Simulation of Pebble Bed Reactors Using Pebble Tracking Transport

The pebble tracking transport (PTT) algorithm offers a high-fidelity deterministic approach for neutron transport for pebble bed reactors (PBRs). This approach requires the mesh for the active-core region to consist exclusively of tetrahedral elements, where each node in the pebble-packing region represents a pebble centroid. This paper investigates the application of PTT for full-scale PBRs, considering both the isothermal and the temperature-dependent core conditions. Macroscopic cross sections are generated using Serpent 2 full-core eigenvalue simulations where pebbles are grouped into disjoint subsets using machine learning. To minimize the need for individual cross-section sets for each pebble in the core, K-means clustering is used to group pebbles by temperature and neutronic environment parameters. Here, we compare the multiplication factor and power rate distributions between PTT simulations using the Griffin reactor physics software and reference solutions from Serpent 2. Our analysis shows that a full-core, high-fidelity PTT calculation produces accurate results with minimal local (pebblewise) errors. Additionally, timing results indicate that PTT simulations converge rapidly on modern supercomputing platforms.

Griffin↗

Systematic characterization of unknown compounds via dimensionality reduction of time series

Analysis of ambient aerosols provides valuable insight into particle sources and formation chemistry. However, due to the complexity of atmospheric data and the dynamic nature of aerosol composition, a substantial fraction of data often become discarded by conventional analysis methods. Furthermore, a large fraction of chemical species within those data are unidentifiable due to a lack of matching spectral information, resulting in suboptimal characterization of chemical composition. Previous work has demonstrated techniques for cataloging analytes in a chromatographic dataset by deconvolution of mass spectra, but integration of these analytes throughout a large dataset remains time consuming. Here, we present a method to automatically identify an ion for quantitation for single-ion chromatogram based peak fitting and integration, enabling comprehensive integration of analytes with minimal user interaction. The resulting time series are clustered with a machine-learning based dimensionality reduction technique to systematically investigate the underlying characteristics of the categorized analytes and gain new insights into the chemical composition and physicochemical properties of the unidentifiable analytes. We apply these methods to existing atmospheric datasets collected in Manacapuru, Brazil during the GoAmazon2014/5 campaign to identify new analytes and interpret their variability and transformations in the atmosphere. The analysis results generate 408 time series from cataloged analytes of interest, and the clustering of those time series with spherical k-means results in 8 distinct clusters. We find the analytes form clusters based on their distinct physicochemical properties, demonstrating the method’s ability to systematically identify and selectively filter contaminants and instrumental analytes and characterize the unidentifiable analytes.

54 ENVIRONMENTAL SCIENCES↗

A weather pattern responsible for increasing wildfires in the western United States

Abstract The western United States (U.S.) has been experiencing more severe wildfires, in part due to climate change, but the underlying synoptic patterns and their modulation in driving fire weather is unclear. Here we investigated the relationship between weather regimes (WRs) and fire weather indices, specifically vapor pressure deficit (VPD) and the Canadian Forest Fire Weather Index. By identifying five singular WRs using k-means clustering, we found that a particular regime (WR-2), one characterized by a distinct tripolar wave train pattern over the continental U.S., has exhibited an increased frequency since 1980. The ascribed WR-2 regime was found to be mainly responsible for rising trends in the fire weather indices, especially VPD. Further, the average fire indices of the WR-2 regime played a more important role than the frequency in shaping the rising trends in the fire weather indices. The increased frequency of the WR-2 WR was mainly attributed to anthropogenic forcing and, the year-to-year variation of the frequency was associated with sea surface temperature anomalies over the subtropical eastern Pacific. Human-induced climate change might have furthered the exacerbation of wildfire danger in the western U.S. by modulating the behaviors of WRs and fire weather indices.

Zhang, Wei (ORCID:0000000221698749)↗

Topological data analysis of task-based fMRI data from experiments on schizophrenia

We use methods from computational algebraic topology to study functional brain networks, in which nodes represent brain regions and weighted edges represent similarity of fMRI time series from each region. With these tools, which allow one to characterize topological invariants such as loops in high-dimensional data, we are able to gain understanding into low-dimensional structures in networks in a way that complements traditional approaches based on pairwise interactions. In the present paper, we analyze networks constructed from task-based fMRI data from schizophrenia patients, healthy controls, and healthy siblings of schizophrenia patients using persistent homology, which allows us to explore the persistence of topological structures such as loops at different scales in the networks. We use persistence landscapes, persistence images, and Betti curves to create output summaries from our persistent-homology calculations, and we study the persistence landscapes and images using k-means clustering and community detection. Based on our analysis of persistence landscapes, we find that the members of the sibling cohort have topological features (specifically, their 1-dimensional loops) that are distinct from the other two cohorts. From the persistence images, we are able to distinguish all three subject groups and to determine the brain regions in the loops (with four or more edges) that allow us to make these distinctions.

60 APPLIED LIFE SCIENCES↗

Enhancing transfer learning in angle-resolved photoemission spectroscopy (ARPES) with spatially-aware representations via graph convolution

A recent application of machine learning has been to spatially-resolved angle-resolved photoemission spectroscopy (ARPES). Here we advance the state-of-the-art by applying representational learning to transform ARPES data into an embedding space of a pre-trained self-supervised learning model, thus enhancing the pipeline that improves the bandstructure classification and domain assignment/segmentation performance compared to a k-means clustering method. In the current iteration, the real-space information is entered into the domain assignment through the graph convolution method, which improves the transfer learning performance of the original self-supervised model. Lastly, an unsupervised automated tool is developed that incorporates these techniques to enable automatic domain assignment.

ARPES↗

Deep unsupervised learning using spike-timing-dependent plasticity

Abstract Spike-timing-dependent plasticity (STDP) is an unsupervised learning mechanism for spiking neural networks that has received significant attention from the neuromorphic hardware community. However, scaling such local learning techniques to deeper networks and large-scale tasks has remained elusive. In this work, we investigate a Deep-STDP framework where a rate-based convolutional network, that can be deployed in a neuromorphic setting, is trained in tandem with pseudo-labels generated by the STDP clustering process on the network outputs. We achieve 24.56% higher accuracy and 3.5 × faster convergence speed at iso-accuracy on a 10-class subset of the Tiny ImageNet dataset in contrast to a k -means clustering approach.

Lu, Sen↗

The drivers and predictability of wildfire re-burns in the western United States (US)

Evidence is mounting that the effectiveness of using prescribed burns as a management tactic may be diminishing due to the higher incidence of wildfire re-burns. The development of predictive models of re-burns is thus essential to better understand their primary drivers so that forest management practices can be updated to account for these events. First, we assess the potential for human activity as a driver of re-burns by evaluating re-burn trends both within and outside of the wildland–urban interface (WUI) of the western US. Next, we investigate the predictability of re-burns through the application of both random forest and the explanatory machine learning non-negative matrix factorization using k-means clustering (NMFk) algorithms to predict re-burn occurrence over California based on a number of climate factors. Our findings indicate that while most states showed increasing trends within the WUI when trends were conducted over longer moving windows (e.g. 20 years), California was the only state where the rate of increase was consistently higher in the WUI, indicating a stronger potential for human activity as a driver in that location. Furthermore, we find model performance was found to be robust over most of California (Testing F1 scores = 0.688), although results were highly variable based on EPA level III Ecoregion (F1 scores = 0.0–0.778). Insights provided from this study will lead to a better understanding of climate and human activity drivers of re-burns and how these vary at broad spatial scales so that improvements in forest management practices can be tuned according to the level of change that is expected for a given region.

54 ENVIRONMENTAL SCIENCES↗

Nanoscale elemental and morphological imaging of nitrogen-fixing cyanobacteria

Nitrogen-fixing cyanobacteria bind atmospheric nitrogen and carbon dioxide using sunlight. This experimental study focused on a laboratory-based model system, Anabaena sp., in nitrogen-depleted culture. When combined nitrogen is scarce, the filamentous prokaryotes reconcile photosynthesis and nitrogen fixation by cellular differentiation into heterocysts. To better understand the influence of micronutrients on cellular function, 2D and 3D synchrotron X-ray fluorescence mappings were acquired from whole biological cells in their frozen-hydrated state at the Bionanoprobe, Advanced Photon Source. To study elemental homeostasis within these chain-like organisms, biologically relevant elements were mapped using X-ray fluorescence spectroscopy and energy-dispersive X-ray microanalysis. Higher levels of cytosolic K + , Ca 2+ , and Fe 2+ were measured in the heterocyst than in adjacent vegetative cells, supporting the notion of elevated micronutrient demand. P-rich clusters, identified as polyphosphate bodies involved in nutrient storage, metal detoxification, and osmotic regulation, were consistently co-localized with K + and occasionally sequestered Mg 2+ , Ca 2+ , Fe 2+ , and Mn 2+ ions. Machine-learning-based k-mean clustering revealed that P/K clusters were associated with either Fe or Ca, with Fe and Ca clusters also occurring individually. In accordance with XRF nanotomography, distinct P/K-containing clusters close to the cellular envelope were surrounded by larger Ca-rich clusters. The transition metal Fe, which is a part of nitrogenase enzyme, was detected as irregularly shaped clusters. The elemental composition and cellular morphology of diazotrophic Anabaena sp. was visualized by multimodal imaging using atomic force microscopy, scanning electron microscopy, and fluorescence microscopy. This paper discusses the first experimental results obtained with a combined in-line optical and X-ray fluorescence microscope at the Bionanoprobe.

Anabaena sp↗

Understanding nanoscale structural distortions in Pb(Zr 0.2 Ti 0.8 )O 3 by utilizing X-ray nanodiffraction and clustering algorithm analysis

Hard X-ray nanodiffraction provides a unique nondestructive technique to quantify local strain and structural inhomogeneities at nanometer length scales. However, sample mosaicity and phase separation can result in a complex diffraction pattern that can make it challenging to quantify nanoscale structural distortions. In this work, a k-means clustering algorithm was utilized to identify local maxima of intensity by partitioning diffraction data in a three-dimensional feature space of detector coordinates and intensity. This technique has been applied to X-ray nanodiffraction measurements of a patterned ferroelectric PbZr 0.2 Ti 0.8 O 3 sample. The analysis reveals the presence of two phases in the sample with different lattice parameters. A highly heterogeneous distribution of lattice parameters with a variation of 0.02 Å was also observed within one ferroelectric domain. This approach provides a nanoscale survey of subtle structural distortions as well as phase separation in ferroelectric domains in a patterned sample.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Partitioning of Large-Scale Power Electronics-Based Power Systems for Small-Signal Stability Analysis

The nodal admittance matrix (NAM)-based approach is suitable for analyzing the small-signal stability of large-scale power electronics-based power systems (PEPSs) as it preserves the system structure by utilizing the admittance matrix. Previously, NAM-based area partition has been proposed, which divides the system into various subareas and interconnections for easier analysis of the low-dimension matrix compared to the entire system-based high-dimension matrix. However, no partition algorithm has been presented for the NAM-based area partition method. This paper focuses on implementing the spectral partitioning algorithm for partitioning large-scale PEPSs into a low-dimension matrix to reduce the computation complexity of the analysis. These spectral components facilitate data transformation into a new space, enabling the application of traditional clustering methods like k-means. To evaluate the performance of the partitioning method, the subareas and interconnections obtained from the spectral clustering algorithm are incorporated into the NAM-based area partition method for a large system with 140 buses. The computational times of the original method, where the NAM-based criterion is directly applied to the entire system, are compared with those of the NAM-based partition method in MATLAB. PSCAD simulations of the whole system and the obtained subareas are conducted to validate the effectiveness of the proposed algorithm.

Nupur, Nupur↗

Transfer Learning Trained LSTM Models for Household Load Profile Forecasting

Grid edge renewable energy resources, such as rooftop solar photovoltaics, closely interact with consumer load profiles. Therefore, forecasting future electricity demand, ideally at the individual household level, is indispensable. In this paper, we present a transfer learning enhanced household load profile forecasting method. First, we tune a long short-term memory forecasting model to perform day-ahead prediction of household electricity load profiles. Then we improve these individualized models using transfer learning, and we use k-means clustering to create optimal source data sets. We find average improvements of 4.38% (largest improvement of 10.71%) when the entire data set was used to train the source model and 2.45% (largest improvement of 11.57%) in the mean absolute error when households were first clustered and used to train separate source models for each cluster. We find that transfer learning with clustered data can effectively boost the forecasting performance of the LSTM models. We use realistic household power measurements for 148 real residential households in Austin, Texas.

deep learning↗

Joint Management and Optimization of Residential Natural Gas and Electricity Distribution Networks Coupled via Fuel Cells

The attractive features of natural gas as well as the growing electric power demand worldwide have created increasing interest in natural-gas-based distributed generation applications for electric distribution networks. Here, this paper investigates the interdependency between a residential natural gas network and an electric distribution network that are linked together via fuel cells. The modeling of the natural gas network is introduced first, and then the algorithm for gas flow study is presented. The optimal placement and sizing of fuel cell based distributed generation systems are formulated to minimize the losses in both the natural gas network and the electric distribution grid, subject to the constraints imposed by both networks. In addition, a probabilistic model for both gas and electricity demands is developed based on historical electricity and natural gas demand data. A K-means clustering method is used to determine the hourly load states to solve the joint probabilistic optimization problem. Simulation studies are carried out on an integrated system consisting of the IEEE 69-bus distribution network and a radial 27-node natural gas network to verify the developed optimization model and the proposed method.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A Hydrogen Load Modeling Method for Integrated Hydrogen Energy System Planning

The integrated hydrogen energy system incorporates hydrogen energy into the power grid, which has been recognized as a promising option for reaching a 100% renewable electricity supply. It can make a profit because the hydrogen produced can be sold as fuel or used to generate electricity for grid services. In this paper, we develop a planning model for the integrated hydrogen energy system that considers the uncertainty of the load demand, the renewable energy generation, and the market prices. To calculate the hydrogen load, we simulate the refueling operations at a hydrogen fueling station over the course of one day and generate representative load profiles with K-means clustering. Moreover, the long-term profitability of the integrated system under both current and future conditions is validated in 10-year planning results.

grid service↗

Synchronization and Rebound Effects in Residential Loads

Increasing fuel prices and capacity investment deferral place an increasing demand for peak reduction from distribution level systems. Residential and commercial devices, such as HVAC systems and water heaters, are increasingly involved in load control programs, and their use may generate synchronization and rebound effects, such as artificial peaks caused by device optimization. While there have been concerns over device synchronization, few studies quantify the extent of this effect with numerical values. In this study, we attempt to investigate whether control efforts result in device synchronization or rebound effects. We focus on three clustering methods – Ward’s clustering, Euclidean K-means, and Density-based spatial clustering of applications with noise – to evaluate the extent of synchronization of a fleet of water heaters and HVAC systems in Atlanta, Georgia. Our findings show that synchronization and rebound effects are present in the neighborhood’s water heaters, but none were found in the HVAC systems. Further, high usage water heaters are more susceptible to synchronization and rebound effects.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Assessing Machine Learning as a Tool to Explain Variance in Deployed Photovoltaic (PV) System Degradation

Degradation remains a large uncertainty in forecasting production for PV plants, creating significant risk for developers and financiers. This study aims to quantify the distribution and drivers of degradation across 10,000 PV systems deployed for distributed or utility generation by training a machine learning model to predict year-over-year degradation rates from metadata characteristics. A combination of K-Means clustering and random forest regressor were found to associate multiple metadata features as potential drivers of degradation, including module characteristics, system design, and climate features. From this, it is inferred that if machine learning is able to find complex patterns between metadata features and system performance loss, such methods can be employed to help developers and financiers make data-informed decisions when estimating long-term energy production forecasts in financial models.

Dunn, Jimmy C.↗

Representative Period Selection for Robust Capacity Expansion Planning in Low-carbon Grids

With the increasing urgency to decarbonize power systems, while mitigating extreme events, capacity expansion models can play a vital role in reliably planning the expansion of power systems and facilitating the integration of renewable energy sources. Optimizing capacity expansion generally involves selecting surrogate representative days from forecasts of load and the generation profiles of variable renewable energy resources. To properly select those representative days, we propose a novel input-based approach in combination with the k-means clustering algorithm that utilize three unique operational inputs: load shedding, renewable curtailment, and transmission congestion. The proposed method allows for more robust and cost-effective capacity planning. The method is validated using a capacity expansion model and a production cost model based on California Independent System Operator (CAISO)'s decarbonization goals, and results in reduced costs and drastically lower load shedding.

Anderson, Osten P.↗

Machine Learning Based Resilience Testing of an Address Randomization Cyber Defense

Moving target defenses (MTDs) are widely used as an active defense strategy for thwarting cyberattacks on cyber-physical systems by increasing diversity of software and network paths. Recently, machine Learning (ML) and deep Learning (DL) models have been demonstrated to defeat some of the cyber defenses by learning attack detection patterns and defense strategies. It raises concerns about the susceptibility of MTD to ML and DL methods. Here, in this article, we analyze the effectiveness of ML and DL models when it comes to deciphering MTD methods and ultimately evade MTD-based protections in real-time systems. Specifically, we consider a MTD algorithm that periodically randomizes address assignments within the MIL-STD-1553 protocol—a military standard serial data bus. Two ML and DL-based tasks are performed on MIL-STD-1553 protocol to measure the effectiveness of the learning models in deciphering the MTD algorithm: 1) determining whether there is an address assignments change i.e., whether the given system employs a MTD protocol and if it does 2) predicting the future address assignments. The supervised learning models (random forest and k-nearest neighbors) effectively detected the address assignment changes and classified whether the given system is equipped with a specified MTD protocol. On the other hand, the unsupervised learning model (K-means) was significantly less effective. The DL model (long short-term memory) was able to predict the future addresses with varied effectiveness based on MTD algorithm's settings.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗