Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “clustering statistics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Mitigating imaging systematics for DESI 2024 emission Line Galaxies and beyond

Emission Line Galaxies (ELGs) are one of the main tracers that the Dark Energy Spectroscopic Instrument (DESI) uses to probe the universe. However, they are afflicted by strong spurious correlations between target density and observing conditions known as imaging systematics. In this paper, we present the imaging systematics mitigation applied to the DESI Data Release 1 (DR1) large-scale structure catalogs used in the DESI 2024 cosmological analyses. We also explore extensions of the fiducial treatment. This includes a combined approach, through forward image simulations (Obiwan) in conjunction with neural network-based regression, to obtain an angular selection function that mitigates the imaging systematics observed in the DESI DR1 ELGs target density. We further derive a line of sight selection function from the forward model that removes the strong redshift dependence between imaging systematics and low redshift ELGs. Combining both angular and redshift-dependent systematics, we construct a three-dimensional selection function and assess the impact of all selection functions on clustering statistics. We quantify differences between these extended treatments and the fiducial treatment in terms of the measured 2-point statistics. We find that the results are generally consistent with the fiducial treatment and conclude that the differences are far less than the imaging systematics uncertainty included in DESI 2024 full-shape measurements. We extend our investigation to the ELGs at 0.6 < z < 0.8, i.e., beyond the redshift range (0.8 < z < 1.6) adopted for the DESI clustering catalog, and demonstrate that determining the full three-dimensional selection function is necessary in this redshift range. Our tests showed that all changes are consistent with statistical noise for BAO analyses indicating they are robust to even severe imaging systematics. Specific tests for the full-shape analysis will be presented in a companion paper.

79 ASTRONOMY AND ASTROPHYSICS↗

Cosmological constraints from density-split clustering in the BOSS CMASS galaxy sample

We present a clustering analysis of the BOSS DR12 CMASS galaxy sample, combining measurements of the galaxy two-point correlation function and density-split clustering down to a scale of $1 \, h^{-1}\, \text{Mpc}$. Our theoretical framework is based on emulators trained on high-fidelity mock galaxy catalogues that forward model the cosmological dependence of the clustering statistics within an extended-ΛCDM framework, including redshift-space and Alcock–Paczynski distortions. Our base-ΛCDM analysis finds ω cdm = 0.1201 ± 0.0022, σ 8 = 0.792 ± 0.034, and n s = 0.970 ± 0.018, corresponding to fσ 8 = 0.462 ± 0.020 at z ≈ 0.525, which is in agreement with Planck 2018 predictions and various clustering studies in the literature. We test single-parameter extensions to base-ΛCDM, varying the running of the spectral index, the dark energy equation of state, and the density of mass-less relic neutrinos, finding no compelling evidence for deviations from the base model. We model the galaxy–halo connection using a halo occupation distribution framework, finding signatures of environment-based assembly bias in the data. We validate our pipeline against mock catalogues that match the clustering and selection properties of CMASS, showing that we can recover unbiased cosmological constraints even with a volume 84 times larger than the one used in this study.

79 ASTRONOMY AND ASTROPHYSICS↗

BASS. XXXVI. Constraining the Local Supermassive Black Hole–Halo Connection with BASS DR2 AGNs

We investigate the connection between supermassive black holes (SMBHs) and their host dark matter halos in the local universe using the clustering statistics and luminosity function of active galactic nuclei (AGNs) from the Swift/BAT AGN Spectroscopic Survey (BASS DR2). By forward-modeling AGN activity into snapshot halo catalogs from N -body simulations, we test a scenario in which SMBH mass correlates with dark matter (sub)halo mass for fixed stellar mass. We compare this to a model absent of this correlation, where stellar mass alone determines the SMBH mass. We find that while both simple models are able to largely reproduce the abundance and overall clustering of AGNs, the model in which black hole mass is tightly correlated with halo mass is preferred by the data by 1.8 σ . When including an independent measurement on the black hole mass–halo mass correlation, this model is preferred by 4.6 σ . We show that the clustering trends with black hole mass can further break the degeneracies between the two scenarios and that our preferred model reproduces the measured clustering differences on one-halo scales between large and small black hole masses. These results indicate that the halo binding energy is fundamentally connected to the growth of SMBHs.

79 ASTRONOMY AND ASTROPHYSICS↗

Failure Mode Identification Through Clustering Analysis

Research has shown that nearly 80% of the costs and problems are created in product development and that cost and quality are essentially designed into products in the conceptual stage. Currently, failure identification procedures (such as FMEA (Failure Modes and Effects Analysis), FMECA (Failure Modes, Effects and Criticality Analysis) and FTA (Fault Tree Analysis)) and design of experiments are being used for quality control and for the detection of potential failure modes during the detail design stage or post-product launch. Though all of these methods have their own advantages, they do not give information as to what are the predominant failures that a designer should focus on while designing a product. This work uses a functional approach to identify failure modes, which hypothesizes that similarities exist between different failure modes based on the functionality of the product/component. In this paper, a statistical clustering procedure is proposed to retrieve information on the set of predominant failures that a function experiences. The various stages of the methodology are illustrated using a hypothetical design example.

Arunajadai, Srikesh G.↗

Window convolution of the galaxy clustering bispectrum

In galaxy survey analysis, the observed clustering statistics do not directly match theoretical predictions but rather have been processed by a window function that arises from the survey geometry including the sky footprint, redshift-dependent background number density and systematic weights. While window convolution of the power spectrum is well studied, for the bispectrum with a larger number of degrees of freedom, it poses a significant numerical and computational challenge. In this work, we consider the effect of the survey window in the tripolar spherical harmonic decomposition of the bispectrum and lay down a formal procedure for their convolution via a series expansion of configuration-space three-point correlation functions, which was first proposed by Sugiyama et al. (2019). We then provide a linear algebra formulation of the full window convolution, where an unwindowed bispectrum model vector can be directly premultiplied by a window matrix specific to each survey geometry. To validate the pipeline, we focus on the Dark Energy Spectroscopic Instrument (DESI) Data Release 1 (DR1) luminous red galaxy (LRG) sample in the South Galactic Cap (SGC) in the redshift bin 0.4 ≤ z ≤ 0.6. We first perform convergence checks on the measurement of the window function from discrete random catalogues, and then investigate the convergence of the window convolution series expansion truncated at a finite of number of terms as well as the performance of the window matrix. This work highlights the differences in window convolution between the power spectrum and bispectrum, and provides a streamlined pipeline for the latter for current surveys such as DESI and the Euclid mission.

79 ASTRONOMY AND ASTROPHYSICS↗

The construction of large-scale structure catalogs for the Dark Energy Spectroscopic Instrument

We present the technical details on how large-scale structure (LSS) catalogs are constructed from redshifts measured from spectra observed by the Dark Energy Spectroscopic Instrument (DESI). The LSS catalogs provide the information needed to determine the relative number density of DESI tracers as a function of redshift and celestial coordinates and, e.g., determine clustering statistics. We produce catalogs that are weighted subsamples of the observed data, each matched to a weighted `random' catalog that forms an unclustered sampling of the probability density that DESI could have observed those data at each location. Precise knowledge of the DESI observing history and associated hardware performance allows for a determination of the DESI footprint and the number of times DESI has covered it at sub-arcsecond level precision. This enables the completeness of any DESI sample to be modeled at this same resolution. The pipeline developed to create LSS catalogs has been designed to easily allow robustness tests and enable future improvements. We describe how it allows ongoing work improving the match between galaxy and random catalogs, such as including further information when assigning redshifts to randoms, accounting for fluctuations in target density, accounting for variation in the redshift success rate, and accommodating blinding schemes.

79 ASTRONOMY AND ASTROPHYSICS↗

Precise relative magnitude measurement improves fracture characterization during hydraulic fracturing

SUMMARY Microseismic monitoring is an important technique to obtain detailed knowledge of in-situ fracture size and orientation during stimulation to maximize fluid flow throughout the rock volume and optimize production. Furthermore, considering that the frequency of earthquake magnitudes empirically follows a power law (i.e. Gutenberg–Richter), the accuracy of microseismic event magnitude distributions is potentially crucial for seismic risk management. In this study, we analyse microseismicity observed during four hydraulic fracture treatments of the legacy Cotton Valley experiment in 1997 at the Carthage gas field of East Texas, where fractures were activated at the base of the sand-shale Upper Cotton Valley formation. We perform waveform cross-correlation to detect similar event clusters, measure relative amplitude from aligned waveform pairs with a principal component analysis, then measure precise relative magnitudes. The new magnitudes significantly reduce the deviations between magnitude differences and relative amplitudes of event pairs. This subsequently reduces the magnitude differences between clusters located at different depths. Reduction in magnitude differences between clusters suggests that some attenuation-related biases could be effectively mitigated with relative magnitude measurements. The maximum likelihood method is applied to understand the magnitude frequency distributions and quantify the seismogenic index of the clusters. Statistical analyses with new magnitudes suggest that fractures that are more favourably oriented for shear failure have lower b-value and higher seismogenic index, suggesting higher potential for relatively larger earthquakes, rather than fractures subparallel to maximum horizontal principal stress orientation.

58 GEOSCIENCES↗

Light potentials of photosynthetic energy storage in the field: what limits the ability to use or dissipate rapidly increased light energy?

The responses of plant photosynthesis to rapid fluctuations in environmental conditions are critical for efficient conversion of light energy. These responses are not well-seen laboratory conditions and are difficult to probe in field environments. We demonstrate an open science approach to this problem that combines multifaceted measurements of photosynthesis and environmental conditions, and an unsupervised statistical clustering approach. In a selected set of data on mint (Mentha sp.), we show that ‘light potentials’ for linear electron flow and non-photochemical quenching (NPQ) upon rapid light increases are strongly suppressed in leaves previously exposed to low ambient photosynthetically active radiation (PAR) or low leaf temperatures, factors that can act both independently and cooperatively. Further analyses allowed us to test specific mechanisms. With decreasing leaf temperature or PAR, limitations to photosynthesis during high light fluctuations shifted from rapidly induced NPQ to photosynthetic control of electron flow at the cytochrome b 6 f complex. At low temperatures, high light induced lumen acidification, but did not induce NPQ, leading to accumulation of reduced electron transfer intermediates, probably inducing photodamage, revealing a potential target for improving the efficiency and robustness of photosynthesis. We discuss the implications of the approach for open science efforts to understand and improve crop productivity.

59 BASIC BIOLOGICAL SCIENCES↗

An Assessment of Non-Powered Dam Hydropower Development Opportunities in the United States

Retrofitting non-powered dams (NPDs) to add hydropower offers multiple benefits over traditional, new hydropower development. Since much of the civil works infrastructure already exists, many retrofits require minimal new construction and can leverage existing water conveyances and discharge for generation. Although NPDs vary considerably in terms of their defining characteristics, identifying similarities helps support targeted investment and find opportunities to develop solutions for common challenges. With the publication of NPD-related data in 2022 (Hansen et al. 2022b), the US Department of Energy and relevant stakeholders are equipped with key information across roughly 89,000 US NPDs. These data represent the best-available information to date and extend beyond previous assessments of potential power capacity by describing design, operational, socioeconomic, environmental aspects of NPDs (Hadjerioua et al. 2012). To capitalize on these efforts to improve breadth and depth of NPD data access, this study uses a data-driven approach to assess NPD hydropower development opportunities in the United States. To help describe project feasibility drivers for NPDs, recent NPD retrofits were reviewed to examine variability with respect to a variety of characteristics. Based on the assessment of recent retrofits and data availability, several attributes were selected to characterize the remaining NPD population: (1) owner type, (2) 30% exceedance flow (a common design flow that describes the flow that is exceeded by 30% of the flow in the record), (3) hydraulic head, and (4) maximum reservoir storage. These four characteristics were used as inputs to a statistical clustering analysis (a common technique for grouping individuals of a population based on similarities or how closely associated individuals are to one another) of a set of 2,709 NPDs with at least 100 kW potential capacity and recently retrofit dams with available data. With the large dataset of NPDs, the clusters help describe how the population breaks down into different types (i.e., how many types of dams there are and what portion of the population belongs to each cluster).

13 HYDRO ENERGY↗

Robust Dark Energy Constraints with the Dark Energy Spectroscopic Survey (Final Technical Report)

This project developed and applied advanced theoretical, computational, and data-analysis methodologies to extract robust and precise cosmological constraints from the Dark Energy Spectroscopic Instrument (DESI). The work focused on maximizing the scientific return of DESI through optimized survey strategy, novel higher-order clustering statistics, improved modeling of small-scale structure, and rigorous mitigation of observational systematics. Over the award period, the project made substantial contributions to DESI science planning, produced new methods for bispectrum and three-point correlation function analyses, advanced constraints on primordial non-Gaussianity, and delivered widely used software tools. The project also played a major role in training graduate students and a postdoctoral researcher who contributed directly to DESI key projects. The results have significantly enhanced the cosmological reach of DESI and provide a strong foundation for future surveys such as DESI-II and Stage-V experiments.

79 ASTRONOMY AND ASTROPHYSICS↗

Metabolome patterns identify active dechlorination in bioaugmentation consortium SDC-9™

Ultra-high performance liquid chromatography–high-resolution mass spectrometry (UPHLC–HRMS) is used to discover and monitor single or sets of biomarkers informing about metabolic processes of interest. The technique can detect 1000’s of molecules (i.e., metabolites) in a single instrument run and provide a measurement of the global metabolome, which could be a fingerprint of activity. Despite the power of this approach, technical challenges have hindered the effective use of metabolomics to interrogate microbial communities implicated in the removal of priority contaminants. Herein, our efforts to circumvent these challenges and apply this emerging systems biology technique to microbiomes relevant for contaminant biodegradation will be discussed. Chlorinated ethenes impact many contaminated sites, and detoxification can be achieved by organohalide-respiring bacteria, a process currently assessed by quantitative gene-centric tools (e.g., quantitative PCR). This laboratory study monitored the metabolome of the SDC-9™ bioaugmentation consortium during cis-1,2-dichloroethene (cDCE) conversion to vinyl chloride (VC) and nontoxic ethene. Untargeted metabolomics using an UHPLC-Orbitrap mass spectrometer and performed on SDC-9™ cultures at different stages of the reductive dechlorination process detected ~10,000 spectral features per sample arising from water-soluble molecules with both known and unknown structures. Multivariate statistical techniques including partial least squares-discriminate analysis (PLSDA) identified patterns of measurable spectral features (peak patterns) that correlated with dechlorination (in)activity, and ANOVA analyses identified 18 potential biomarkers for this process. Statistical clustering of samples with these 18 features identified dechlorination activity more reliably than clustering of samples based only on chlorinated ethene concentration and Dhc 16S rRNA gene abundance data, highlighting the potential value of metabolomic workflows as an innovative site assessment and bioremediation monitoring tool.

environmental monitoring↗

One Galaxy Sample to Rule Them All: Halo Occupation Distribution Modeling of DES Year 3 Source Galaxies

Abstract For the joint analysis of second-order weak-lensing and galaxy clustering statistics, so-called 3 × 2 analyses, the selection and characterization of optimal galaxy samples is a major area of research. One promising choice is to use the same galaxy sample as lenses and sources, which reduces the systematics parameter space that describes the uncertainties related to galaxy samples. Such a “lens-equal-source” analysis significantly improves the self-calibration of photo- z systematics, leading to improved cosmological constraints. With the aim of enabling a lens-equal-source analysis on small scales, we investigate the halo–galaxy connection of DES Year 3 source galaxies. We develop a technique to construct mock source galaxy populations by matching COSMOS/UltraVISTA photometry to U niverse M achine galaxies. These mocks predict a source halo occupation distribution (HOD) that exhibits significant redshift evolution, nontrivial central incompleteness, and galaxy assembly bias. We produce multiple realizations of mock source galaxies drawn from the U niverse M achine posterior, with added uncertainties in the measured Dark Energy Survey photometry and galaxy shapes. We fit a modified HOD formalism to these realizations to produce priors on the galaxy–halo connection for cosmological analyses. We additionally train an emulator that predicts this HOD to ∼2% accuracy from redshift z = 0.1−1.3 that models the dependence of this HOD on (1) observational uncertainties in galaxy size and photometry and (2) uncertainties in the U niverse M achine predictions.

Salcedo, Andrés N. (ORCID:000000031420527X)↗

Unsupervised classification of remote multispectral sensing data

The new unsupervised classification technique for classifying multispectral remote sensing data which can be either from the multispectral scanner or digitized color-separation aerial photographs consists of two parts: (a) a sequential statistical clustering which is a one-pass sequential variance analysis and (b) a generalized K-means clustering. In this composite clustering technique, the output of (a) is a set of initial clusters which are input to (b) for further improvement by an iterative scheme. Applications of the technique using an IBM-7094 computer on multispectral data sets over Purdue's Flight Line C-1 and the Yellowstone National Park test site have been accomplished. Comparisons between the classification maps by the unsupervised technique and the supervised maximum liklihood technique indicate that the classification accuracies are in agreement.

Su, M. Y.↗

Nineteen hundred seventy three significant accomplishments

Data collected by the Skylab remote sensing satellites was used to develop applications techniques and to combine automatic data classification with statistical clustering methods. Continuing research was concentrated in the correlation and registration of data products and in the definition of the atmospheric effects on remote sensing. The causes of errors encountered in the automated classification of agricultural data are identified. Other applications in forestry, geography, environmental geology, and land use are discussed.

Source record↗

Analysis of the Tanana River Basin using LANDSAT data

Digital image classification techniques were used to classify land cover/resource information in the Tanana River Basin of Alaska. Portions of four scenes of LANDSAT digital data were analyzed using computer systems at Ames Research Center in an unsupervised approach to derive cluster statistics. The spectral classes were identified using the IDIMS display and color infrared photography. Classification errors were corrected using stratification procedures. The classification scheme resulted in the following eleven categories; sedimented/shallow water, clear/deep water, coniferous forest, mixed forest, deciduous forest, shrub and grass, bog, alpine tundra, barrens, snow and ice, and cultural features. Color coded maps and acreage summaries of the major land cover categories were generated for selected USGS quadrangles (1:250,000) which lie within the drainage basin. The project was completed within six months.

Morrissey, L. A.↗

MEASURE: An integrated data-analysis and model identification facility

The first phase of the development of MEASURE, an integrated data analysis and model identification facility is described. The facility takes system activity data as input and produces as output representative behavioral models of the system in near real time. In addition a wide range of statistical characteristics of the measured system are also available. The usage of the system is illustrated on data collected via software instrumentation of a network of SUN workstations at the University of Illinois. Initially, statistical clustering is used to identify high density regions of resource-usage in a given environment. The identified regions form the states for building a state-transition model to evaluate system and program performance in real time. The model is then solved to obtain useful parameters such as the response-time distribution and the mean waiting time in each state. A graphical interface which displays the identified models and their characteristics (with real time updates) was also developed. The results provide an understanding of the resource-usage in the system under various workload conditions. This work is targeted for a testbed of UNIX workstations with the initial phase ported to SUN workstations on the NASA, Ames Research Center Advanced Automation Testbed.

Singh, Jaidip↗

A Fast Implementation of the ISOCLUS Algorithm

Unsupervised clustering is a fundamental building block in numerous image processing applications. One of the most popular and widely used clustering schemes for remote sensing applications is the ISOCLUS algorithm, which is based on the ISODATA method. The algorithm is given a set of n data points in d-dimensional space, an integer k indicating the initial number of clusters, and a number of additional parameters. The general goal is to compute the coordinates of a set of cluster centers in d-space, such that those centers minimize the mean squared distance from each data point to its nearest center. This clustering algorithm is similar to another well-known clustering method, called k-means. One significant feature of ISOCLUS over k-means is that the actual number of clusters reported might be fewer or more than the number supplied as part of the input. The algorithm uses different heuristics to determine whether to merge lor split clusters. As ISOCLUS can run very slowly, particularly on large data sets, there has been a growing .interest in the remote sensing community in computing it efficiently. We have developed a faster implementation of the ISOCLUS algorithm. Our improvement is based on a recent acceleration to the k-means algorithm of Kanungo, et al. They showed that, by using a kd-tree data structure for storing the data, it is possible to reduce the running time of k-means. We have adapted this method for the ISOCLUS algorithm, and we show that it is possible to achieve essentially the same results as ISOCLUS on large data sets, but with significantly lower running times. This adaptation involves computing a number of cluster statistics that are needed for ISOCLUS but not for k-means. Both the k-means and ISOCLUS algorithms are based on iterative schemes, in which nearest neighbors are calculated until some convergence criterion is satisfied. Each iteration requires that the nearest center for each data point be computed. Naively, this requires O(kn) time, where k denotes the current number of centers. Traditional techniques for accelerating nearest neighbor searching involve storing the k centers in a data structure. However, because of the iterative nature of the algorithm, this data structure would need to be rebuilt with each new iteration. Our approach is to store the data points in a kd-tree data structure. The assignment of points to nearest neighbors is carried out by a filtering process, which successively eliminates centers that can not possibly be the nearest neighbor for a given region of space. This algorithm is significantly faster, because large groups of data points can be assigned to their nearest center in a single operation. Preliminary results on a number of real Landsat datasets show that our revised ISOCLUS-like scheme runs about twice as fast.

Memarsadeghi, Nargess↗

A Fast Implementation of the ISOCLUS Algorithm

Unsupervised clustering is a fundamental tool in numerous image processing and remote sensing applications. For example, unsupervised clustering is often used to obtain vegetation maps of an area of interest. This approach is useful when reliable training data are either scarce or expensive, and when relatively little a priori information about the data is available. Unsupervised clustering methods play a significant role in the pursuit of unsupervised classification. One of the most popular and widely used clustering schemes for remote sensing applications is the ISOCLUS algorithm, which is based on the ISODATA method. The algorithm is given a set of n data points (or samples) in d-dimensional space, an integer k indicating the initial number of clusters, and a number of additional parameters. The general goal is to compute a set of cluster centers in d-space. Although there is no specific optimization criterion, the algorithm is similar in spirit to the well known k-means clustering method in which the objective is to minimize the average squared distance of each point to its nearest center, called the average distortion. One significant feature of ISOCLUS over k-means is that clusters may be merged or split, and so the final number of clusters may be different from the number k supplied as part of the input. This algorithm will be described in later in this paper. The ISOCLUS algorithm can run very slowly, particularly on large data sets. Given its wide use in remote sensing, its efficient computation is an important goal. We have developed a fast implementation of the ISOCLUS algorithm. Our improvement is based on a recent acceleration to the k-means algorithm, the filtering algorithm, by Kanungo et al.. They showed that, by storing the data in a kd-tree, it was possible to significantly reduce the running time of k-means. We have adapted this method for the ISOCLUS algorithm. For technical reasons, which are explained later, it is necessary to make a minor modification to the ISOCLUS specification. We provide empirical evidence, on both synthetic and Landsat image data sets, that our algorithm's performance is essentially the same as that of ISOCLUS, but with significantly lower running times. We show that our algorithm runs from 3 to 30 times faster than a straightforward implementation of ISOCLUS. Our adaptation of the filtering algorithm involves the efficient computation of a number of cluster statistics that are needed for ISOCLUS, but not for k-means.

Memarsadeghi, Nargess↗