Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “k mean”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Joint Management and Optimization of Residential Natural Gas and Electricity Distribution Networks Coupled via Fuel Cells

The interesting properties of natural gas as well as the growing electric power demand worldwide have led to increasing attention to natural-gas-based distributed generation applications in electric distribution systems. This paper goes over the interdependency between a residential natural gas network and an electric distribution network that are coupled via fuel cells. The modeling of the gas network is introduced first, and then the algorithm for gas flow study is presented. The optimal placement and sizing of fuel cell based distributed generation systems are formulated to minimize the losses in both the gas and electric distribution networks, subject to their model constraints. In addition to this, in order to capture the probabilistic nature of the optimization problem under study, the K-means clustering algorithm is applied to the gas and electricity demands to determine hourly load states and their corresponding probabilities. Furthermore, simulation studies are carried out on an integrated system consisting of the IEEE 69-bus distribution feeder and a radial 27-node natural gas network to verify the developed optimization model and the proposed method.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Machine Learning-Driven Reliability Estimation of PV Inverters Considering Alert-Ambient Variability

Weather-induced spatio-temporal degradation limits outdoor PV inverter lifetime and reliability, necessitating advanced data analysis. This study employs a top-down, data-driven approach utilizing multiple machine learning (ML) algorithms to estimate inverter reliability in a 1.4 MW PV power plant, considering factors such as irradiance, humidity, temperature, time of day, and weather conditions. An extensive alert dataset from 17 identical inverters, including alert types, propagation, and frequency, reveals significant correlations with environmental factors and inverter output power, enabling the construction of a performance reliability model. Dual-stage supervised-ML models are evaluated for accuracy, with the ‘classification-regression’ model by an artificial neural network (ANN) tested on the averaged “Alert-Ambient” dataset, which is outperformed by ‘clustering-regression’ models using random forest (RF) and K-Nearest Neighbors (KNN) on individual inverter datasets. K-means clustering applies principal component analysis to reduce dimensions, achieving improved accuracy beyond the 80% achieved by ANN on the averaged dataset. Second-stage regression estimates inverter reliability with a mean square error of 0.0195 on the averaged dataset and as low as 0.002 on individual inverter datasets using RF. Furthermore, these findings highlight the method's suitability for estimating PV inverter output reliability under ambient conditions, essential for digital twin development and related applications.

14 SOLAR ENERGY↗

Crash Risks Evaluation of Urban Expressways: A Case Study in Shanghai

We report that proactive traffic safety management systems can reduce crashes by identifying crash precursors, evaluating real-time crash risks, and implementing suitable interventions. The basic prerequisite for developing such a system is to propose a reliable crash risk evaluation model that takes real-time traffic flow data as input. Previous studies have primarily focused on real-time crash prediction using some statistical or machine-learning methods. However, further quantitative evaluation and classification of crash risks have been ignored. In this study, we conduct a systematic crash risk evaluation workflow, including crash risk prediction, crash risk quantification, and crash risk classification. Specifically, the crash risk prediction using an extended logit model is proposed, from which CAS, CSD, UAS, DAS, DTV are identified to be contributing factors of crash risks. Then a crash risk quantification model based on the parameter evaluation of the extended logit model is developed. The crash risks of urban expressways and their spatial-temporal evolution trends are quantified. Finally, the crash risks are classified into high crash risk level, moderate crash risk level, and low crash risk level by the k-means cluster algorithm. Then the threshold boundaries of different crash risk levels are determined. The research results provide a proactive guidance for traffic safety management of urban expressways.

33 ADVANCED PROPULSION SYSTEMS↗

Resilience-Oriented DG Siting and Sizing Considering Stochastic Scenario Reduction

In this paper, a fuel-based distributed generator (DG) allocation strategy is proposed to enhance the distribution system resilience against extreme weather. The long-term planning problem is formulated as a two-stage stochastic mixed-integer programming (SMIP). The first stage is to make decisions of DG siting and sizing under the given budget constraint. In the second stage, a post-extreme-event-restoration (PEER) is employed to minimize the operating cost in an uncertain fault scenario. In particular, this study proposes a method to select the most representative scenarios for the SMIP. First, a Monte Carlo Simulation (MCS) is introduced to generate sufficient scenarios considering random fault locations and load profiles. Then, the number of scenarios is reduced by the K-means clustering algorithm. The advantage of scenario reduction is to make a trade-off between accuracy and computational efficiency. Finally, the SMIP is solved by the progressive hedging algorithm. Here, the case studies of the IEEE 33-bus and 123-bus test systems demonstrate the effectiveness of the proposed algorithm in reducing the expected energy not served (EENS), which is a critical criterion of resilience.

42 ENGINEERING↗

An Artificial-Intelligence and Machine-Learning-Based Methodology to Conduct Seemingly Strain-Controlled Fatigue Test in a Pressurized-Water-Reactor-Test-Loop-Autoclave, While Not Controlling the Strain

In general, low cycle fatigue analysis of pressurized water reactor (PWR) components, requires strain-controlled fatigue test data such as using strain versus life (ε–N) curves. Conducting strain-controlled fatigue tests under in-air conditions is not an issue. However, controlling strain in a PWR-test-loop-autoclave is a challenge, since an extensometer cannot be placed in a narrow autoclave (typically used in a high-temperature-pressure PWR-test-loop). This is due to lack of space inside an autoclave that houses the test specimen. In addition, installing a contact-type extensometer in the path of a high-pressure flow can be a challenge. These difficulties of using an extensometer inside an autoclave led us to use an outside-autoclave displacement sensor which measures the displacement of pull-rod-specimen assembly. However, in our study (based on in-air fatigue test data), we found that a pull-rod-controlled based fatigue test can lead to substantial cyclic hardening/softening resulting in substantially different cyclic strain amplitudes and their rates compared to the desired cyclic strain amplitudes and its rates. In this paper, we propose an Artificial-Intelligence and Machine-Learning based technique such as using k-means clustering technique to improve the pull-rod-control based fatigue test method, such that the gage-area strain amplitude and rates can reasonably be achieved. In support of this, we present the fatigue test results for both 316 SS base and 81/182 dissimilar-metal-weld specimens.

42 ENGINEERING↗

Quantifying and Zoning Urban Heat Island Effects Using Unsupervised Machine Learning

This work explores the Urban Heat Island (UHI) effects in Maricopa County, Arizona, employing a simulation-based approach that combines large-scale building energy modeling with advanced spatial analysis. Utilizing the Automatic Building Energy Modeling (AutoBEM) software suite, we simulated the energy consumption for approximately 1.35 million buildings based on the Model America version 1.0 (MAv1) dataset. Our methodology incorporated spatial analysis at multiple scales, including individual buildings, clusters of zones determined by K-means clustering, and geographical level evaluation based on Zip codes. The results revealed significant variations in energy consumption and heat emissions across different building types and urban zones. High-emission hotspots identified through clustering pointed to areas most contributing to the UHI effects. Zip code-based area analysis further contextualized these findings, offering an urban context-based perspective on emission distribution and informing potential urban energy policies for mitigating UHI effects.

Chowdhury, Shovan [ORNL]↗

A Shift Selection Strategy for Parallel Shift-invert Spectrum Slicing in Symmetric Self-consistent Eigenvalue Computation

The central importance of large-scale eigenvalue problems in scientific computation necessitates the development of massively parallel algorithms for their solution. Recent advances in dense numerical linear algebra have enabled the routine treatment of eigenvalue problems with dimensions on the order of hundreds of thousands on the world’s largest supercomputers. In cases where dense treatments are not feasible, Krylov subspace methods offer an attractive alternative due to the fact that they do not require storage of the problem matrices. However, demonstration of scalability of either of these classes of eigenvalue algorithms on computing architectures capable of expressing massive parallelism is non-trivial due to communication requirements and serial bottlenecks, respectively. In this work, we introduce the SISLICE method: a parallel shift-invert algorithm for the solution of the symmetric self-consistent field (SCF) eigenvalue problem. The SISLICE method drastically reduces the communication requirement of current parallel shift-invert eigenvalue algorithms through various shift selection and migration techniques based on density of states estimation and k-means clustering, respectively. This work demonstrates the robustness and parallel performance of the SISLICE method on a representative set of SCF eigenvalue problems and outlines research directions that will be explored in future work.

97 MATHEMATICS AND COMPUTING↗

Subseasonal Representation and Predictability of North American Weather Regimes Using Cluster Analysis

Abstract This study focuses on assessing the representation and predictability of North American weather regimes, which are persistent large-scale atmospheric patterns, in a set of initialized subseasonal reforecasts created using the Community Earth System Model, version 2 (CESM2). The k -means clustering was used to extract four key North American (10°–70°N, 150°–40°W) weather regimes within ERA5 reanalysis, which were used to interpret CESM2 subseasonal forecast performance. Results show that CESM2 can recreate the climatology of the four main North American weather regimes with skill but exhibits biases during later lead times with overoccurrence of the West Coast high regime and underoccurrence of the Greenland high and Alaskan ridge regimes. Overall, the West Coast high and Pacific trough regimes exhibited higher predictability within CESM2, partly related to El Niño. Despite biases, several reforecasts were skillful and exhibited high predictability during later lead times, which could be partly attributed to skillful representation of the atmosphere from the tropics to extratropics upstream of North America. The high predictability at the subseasonal time scale of these case-study examples was manifested as an “ensemble realignment,” in which most ensemble members agreed on a prediction despite ensemble trajectory dispersion during earlier lead times. Weather regimes were also shown to project distinct temperature and precipitation anomalies across North America that largely agree with observational products. This study further demonstrates that unsupervised learning methods can be used to uncover sources and limits of subseasonal predictability, along with systematic biases present in numerical prediction systems. Significance Statement North American weather regimes are large-scale atmospheric patterns that can persist for several days. Their skillful subseasonal (2 weeks or greater) prediction can provide valuable lead time to prepare for temperature and precipitation anomalies that can stress energy and water resources. The purpose of this study was to assess the climatological representation and subseasonal predictability of four key North American weather regimes using a research subseasonal prediction system and clustering analysis. We found that the Pacific trough and West Coast high regimes exhibited higher predictability than other regimes and that skillful representation of conditions across the tropics and extratropics can increase predictability during later lead times. Future work will quantify causal pathways associated with high predictability.

58 GEOSCIENCES↗

ARMing the Edge: Designing Edge Computing–Capable Machine Learning Algorithms to Target ARM Doppler Lidar Processing

Abstract There is a need for long-term observations of cloud and precipitation fall speeds in validating and improving rainfall forecasts from climate models. To this end, the U.S. Department of Energy Atmospheric Radiation Measurement (ARM) user facility Southern Great Plains (SGP) site at Lamont, Oklahoma, hosts five ARM Doppler lidars that can measure cloud and aerosol properties. In particular, the ARM Doppler lidars record Doppler spectra that contain information about the fall speeds of cloud and precipitation particles. However, due to bandwidth and storage constraints, the Doppler spectra are not routinely stored. This calls for the automation of cloud and rain detection in ARM Doppler lidar data so that the spectral data in clouds can be selectively saved and further analyzed. During the ARMing the Edge field experiment, a Waggle node capable of performing machine learning applications in situ was deployed at the ARM SGP site for this purpose. In this paper, we develop and test four algorithms for the Waggle node to automatically classify ARM Doppler lidar data. We demonstrate that supervised learning using a ResNet50-based classifier will classify 97.6% of the clear-air images and 94.7% of cloudy images correctly, outperforming traditional peak detection methods. We also show that a convolutional autoencoder paired with k -means clustering identifies 10 clusters in the ARM Doppler lidar data. Three clusters correspond to mostly clear conditions with scattered high clouds, and seven others correspond to cloudy conditions with varying cloud-base heights.

54 ENVIRONMENTAL SCIENCES↗

Projecting Future Energy Production from Operating Wind Farms in North America. Part II: Statistical Downscaling

Abstract Capacity factors (CFs) derived from daily expected power at 22 operating wind farms in different regions of North America are used as predictands to train statistical downscaling algorithms using output from ERA5. The statistical downscaling models are then used to make CF projections for a suite of CMIP6 Earth System Models (ESMs). Downscaling is performed using a hybrid statistical approach that employs synoptic types derived using k -means clustering applied to sea level pressure fields with variance corrections applied as a function of the pressure gradient intensity. ESMs exhibit marked variability in terms of the skill with which the frequency of synoptic types and pressure gradients are reproduced relative to ERA5, and that differential skill is used to infer differential credibility in the associated CF projections. Projections of median annual mean CF [P50(CF)] in each 20-yr period from 1980 to 2099 show evidence of declines at most wind farms except in parts of the southern Great Plains, although the magnitude of the changes is strongly dependent on the ESM. For example, P50(CF) in 2080–99 deviate from those in 1980–99 by from −3.1 to +0.2 percentage points in the Northeast. The largest-magnitude declines in P50(CF) ranging from −3.9 to −2 percentage points are projected for the southern West Coast. CF trends exhibit marked seasonality and are strongly linked to changes in the relative intensity of future synoptic patterns, with much less impact from shifts in the occurrence of synoptic types over time. Internal climate modes continue to play a significant role in inducing interannual variability in wind power production, even under high radiative forcing scenarios. Significance Statement We describe how future climate changes may affect wind resources and wind power generation. Near-term changes in projected wind power electricity generation potential at operating wind farms over North America are small, but by the end of the current century electricity production is projected to decrease in many areas but may increase in parts of the southern Great Plains. The amount of change in projected wind power production is a strong function of the Earth system model that is downscaled and also depends on the continued presence of internally forced climate variability. An additional dependence on the amount of greenhouse gas–induced global warming indicates the transition of the energy sector to low-carbon sources may assist in maintaining the abundant U.S. wind resource.

Meteorology & Atmospheric Sciences↗

Summertime Marine Boundary Layer Cloud, Thermodynamic, and Drizzle Morphology over the Eastern North Atlantic: A Four-Year Study

Abstract Summertime remote sensor and in situ data from 2016 to 2019 collected at the ARM Eastern North Atlantic (ENA) Observatory are combined with aircraft measurements from the Aerosol and Cloud Experiments in the Eastern North Atlantic (ACE-ENA) campaign to quantify marine boundary layer (MBL) cloud, thermodynamic, and drizzle morphology in the region. A radar reflectivity–rainfall rate relationship ( Z – R ) is developed from aircraft data and 6-h cloud morphological regimes are identified from ENA data using a k -means algorithm driven by three independent inputs quantifying cloud thickness, drizzle intensity, and cloud field geometric complexity. Four separate MBL structural regimes representing non- or weakly drizzling single-layer stratocumulus, drizzling stratocumulus and cumulus-coupled stratocumulus, deep convection, and broken clouds embedded in northerly flow are identified. Single-layer stratocumulus is indicated when weak subtropical anticyclones are significantly west of the ENA site, and the MBL is cooler and drier than when drizzling and cumulus-coupled stratocumulus and broken clouds are observed. Drizzling and cumulus-coupled stratocumulus clouds are observed on the eastern flank of strong subtropical anticyclones in deep warm moist air masses with wind speeds exceeding 7 m s −1 and strong near-surface wind shear. Broken clouds exhibit strong wind shear near the inversion, while single-layer stratocumulus clouds have lower wind speeds and minimal shear. Net latent heat fluxes in the subcloud layer resulting from a combination of the ocean surface heat flux and evaporating drizzle average near zero over long periods in drizzling and cumulus-coupled stratocumulus. The ECMWF reanalysis version 5 (ERA5) is found to accurately represent single-layer stratocumulus properties, while producing significant discrepancies when drizzling stratocumulus and cumulus-coupled stratocumulus are observed.

54 ENVIRONMENTAL SCIENCES↗

Assessing CESM2 Clouds and Their Response to Climate Change Using Cloud Regimes

Abstract The Community Earth System Model, version 2 (CESM2), has a very high climate sensitivity driven by strong positive cloud feedbacks. To evaluate the simulated clouds in the present climate and characterize their response with climate warming, a clustering approach is applied to three independent satellite cloud products and a set of coupled climate simulations. Using k -means clustering with a Wasserstein distance cost function, a set of typical cloud configurations is derived for the satellite cloud products. Using satellite simulator output, the model clouds are classified into the observed cloud regimes in both current and future climates. The model qualitatively reproduces the observed cloud configurations in the historical simulation using the same time period as the satellite observations, but it struggles to capture the observed heterogeneity of clouds which leads to an overestimation of the frequency of a few preferred cloud regimes. This problem is especially apparent for boundary layer clouds. Those low-level cloud regimes also account for much of the climate response in the late twenty-first century in four shared socioeconomic pathway simulations. The model reduces the frequency of occurrence of these low-cloud regimes, especially in tropical regions under large-scale subsidence, in favor of regimes that have weaker cloud radiative effects.

58 GEOSCIENCES↗

Understanding Observed Precipitation Change and the New Climate Normal from the Perspective of Daily Weather Types in the Southeast United States

Abstract Observed precipitation changes in the Southeast United States (SEUS) are spatially heterogeneous. Most of the inland SEUS and eastern Gulf Coast become drier, and the East Coast north of Charleston, South Carolina, and southern Florida become wetter from the old 30-yr period of 1961–90 to the recent period of 1991–2020. The observed climate change is examined from the perspective of daily weather types (WTs). A k -means clustering analysis has been conducted using daily 850-hPa circulation for 1948–2021. The obtained 10 WTs peak in different seasons, respectively. The frequencies and precipitation intensity of the WTs have been analyzed. A winter WT characterized by a western Appalachian trough (WAT) and a summer WT featuring North Atlantic subtropical high (NASH) have a rising trend of annual frequency from 1948 to 2021. An Appalachian high in the autumn has a decreasing frequency but becomes drier and stronger. Some precipitation intensity change and small location shift have also been observed. The drying up on the eastern Gulf Coast and the inland area of the SEUS is mainly caused by the weakened southwesterly low-level jet (LLJ) on the western flank of the NASH that reduces rain in the spring, the less frequent but stronger and drier Appalachian high in the summer and autumn, and the weaker and more western located Plains trough (PT) in the winter, spring, and autumn. The precipitation increase in the East Coast and southern Florida is majorly due to more frequent, stronger, and rainier troughs along the western Appalachian as well as the East Coast.

Qian, Jian-Hua↗

Two Large-Scale Meteorological Patterns are Associated with Short-Duration Dry Spells in the Northeastern United States

Large-scale meteorological pattern (LSMP)–based analysis is used in a novel way to understand meteorological conditions before and during short-duration dry spells over the northeastern United States. These LSMPs are useful to assess models and select better-performing models for future projections. Dry-spell events are identified from histograms of consecutive dry days below a daily precipitation threshold. Events lasting 12 days or longer, which correspond to ~10% of dry-spell events, are examined. The 500-hPa stream-function anomaly fields for the first 12 days of each event are time averaged, and k -means clustering is applied to isolate the dry-spell-related LSMPs. The first cluster has a strong low pressure anomaly over the Atlantic Ocean, southeast of the region, and is more common in winter and spring. The second cluster has strong high pressure over east-central North America and is most common during autumn. Over the region, both clusters have negative specific humidity anomalies, negative integrated vapor transport from the north, and subsidence associated with a midlatitude jet stream dipole structure that reinforces upper-level convergence. Subsidence is supported by cold-air advection in the first cluster and the location on the east side of the lower-level high pressure in the second cluster. Extratropical cyclone storm tracks are generally shifted southward of the region during the dry spells. Individual events lie on a continuum between two distinct clusters. These clusters have similar local, but different remote, properties. Although dry spells occur with greater frequency during drought months, most dry spells occur during nondrought months. Significance Statement: This study examines the large-scale weather patterns and meteorological conditions associated with dry-spell events lasting at least 2 weeks while affecting the northeastern United States. A statistical approach groups events together on the basis of similar atmospheric features. We find two distinct sets of patterns that we call large-scale meteorological patterns. These patterns reduce moisture, foster localized sinking, and shift the storm track southward along the Atlantic seaboard, all of which reduce precipitation. Besides greater understanding, knowing the meteorological patterns during short-term dryness in the region provides an important tool to assess how well atmospheric models reproduce these specific patterns. More dry spells occur in non-drought months than in drought months, which means that dry spells can occur without preexisting drought conditions.

58 GEOSCIENCES↗

Five-day track forecast skills of WRF model for the western North Pacific tropical cyclones

In this study, the characteristics of simulated tropical cyclones (TCs) over the western North Pacific by a regional model (the WRF Model) are verified. We utilize 12-km horizontal grid spacing, and simulations are integrated for 5 days from model initialization. A total of 125 forecasts are divided into five clusters through the k-means clustering method. The TCs in the cluster 1 and 2 (group 1), which includes many TCs moving northward in the subtropical region, generally have larger track errors than for TCs in cluster 3 and 4 (group 2). The optimal steering vector is used to examine the difference in the track forecast skill between these two groups. The bias in the steering vector between the model and analysis data is found to be more substantial for group 1 TCs than group 2 TCs. The larger steering vector difference for group 1 TCs indicates that environmental fields tend to be poorly simulated in group 1 TC cases. Furthermore, the residual terms, including the storm-scale process, asymmetric convection distribution, or beta-related effect, are also larger for group 1 TCs than group 2 TCs. Therefore, it is probable that the large track forecast error for group 1 TCs is a result of unreasonable simulations of environmental wind fields and residual processes in the midlatitudes.

54 ENVIRONMENTAL SCIENCES↗

Training dataset and results for geothermal exploration artificial intelligence, applied to Brady Hot Springs and Desert Peak

The submission includes the labeled datasets, as ESRI Grid files (.gri, .grd) used for training and classification results for our machine leaning model: - brady_som_output.gri, brady_som_output.grd, brady_som_output.* - desert_som_output.gri, desert_som_output.grd, desert_som_output.* The data corresponds to two sites: Brady Hot Springs and Desert Peak, both located near Fallon, NV. Input layers include: - Geothermal: Labeled data (0: Non-geothermal; 1: Geothermal) - Minerals: Hydrothermal mineral alterations, as a result of spectral analysis using Chalcedony, Kaolinite, Gypsum, Hematite and Epsomite - Temperature: Land surface temperature (% of times a pixel was classified as "Hot" by K-Means) - Faults: Fault density with a 300mradius - Subsidence: PSInSAR results showing subsidence displacement of more than 5mm - Uplift: PSInSAR results showing subsidence displacement of more than 5mm Also, the results of the classification using Brady and Desert Peak to build 2 Convolutional Neural Networks. These were applied to the training site as well as the other site, the results are in GeoTiff format. - brady_classification: Results of classification of the Brady-trained model - desert_classification: Results of classification of the Desert Peak-trained model - b2d_classification: Results of classification of Desert Peak using the Brady-trained model - d2b_classification: Results of classification of Brady using the Desert Peak-trained model

15 GEOTHERMAL ENERGY↗

The importance of accounting for landscape position when investigating grasslands: A multidisciplinary characterisation of a Californian coastal grassland

Data from the characterisation of the Point Reyes Field Site, published in AGU Earth's Future under the title: The importance of accounting for landscape position when investigating grasslands: A multidisciplinary characterisation of a Californian coastal grassland. This paper explored the effect of landscape position on the response of a Californian grassland to seasonal changes. All files are csv files. The EMI data contains 8 csv files with a metadata csv explaining the columns. The dataset also includes soil variables including total concentrations calculated from fused samples, then dissolved and measured on ICP-AES for whole-rock elements and ICP-MS for trace elements. Mineral composition was attained using X-ray diffraction at BL 11-3 at SSRL . Data was then run through the High Score database to characterise different mineral phases. total It also includes a table with bulk soil characteristics such as soil pH, cation exchange capacity, and soil textural data. Data from Teros 12 Meter soil moisture, electrical conductivity and temperature sensors are presented in SMS Csv file. While the WL bottom and top files contain data from Piezometers measuring the ground water table. We have included a csv file that contains soil CO2 efflux data from Feb 2021-Oct 2021 in the Point Reyes Grassland Experiment We have included the spatially orientated (easting northing) remotely sensed datasets that were used in the K-means clustering analysis conducted on our site with electrical conductivity, normalised difference vegetation index, elevation, slope, solar radiation, topographic position and wetness index, and a clustering score. Finally there is a list of all the identified grassland species at the site.For more information on flux data, please contact the corresponding author.

54 ENVIRONMENTAL SCIENCES↗