An application of statistics in mixture of exponential distributions
Order statistics in application to exponential distributions
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Order statistics in application to exponential distributions
Estimation in mixtures of Poisson and mixtures of exponential distributions
Simulation tests were carried out to compare the power of the Kolmogoroff-Smirnoff and Z tests for the exponential distribution against a wide range of alternative distributions. The results indicate that both tests should be used for applications for which detailed knowledge regarding the possible classes of alternative distributions is lacking.
Explore the source record for details and available documents.
Petri nets augmented with timing specifications gained a wide acceptance in the area of performance and reliability evaluation of complex systems exhibiting concurrency, synchronization, and conflicts. The state space of time-extended Petri nets is mapped onto its basic underlying stochastic process, which can be shown to be Markovian under the assumption of exponentially distributed firing times. The integration of exponentially and non-exponentially distributed timing is still one of the major problems for the analysis and was first attacked for continuous time Petri nets at the cost of structural or analytical restrictions. We propose a discrete deterministic and stochastic Petri net (DDSPN) formalism with no imposed structural or analytical restrictions where transitions can fire either in zero time or according to arbitrary firing times that can be represented as the time to absorption in a finite absorbing discrete time Markov chain (DTMC). Exponentially distributed firing times are then approximated arbitrarily well by geometric distributions. Deterministic firing times are a special case of the geometric distribution. The underlying stochastic process of a DDSPN is then also a DTMC, from which the transient and stationary solution can be obtained by standard techniques. A comprehensive algorithm and some state space reduction techniques for the analysis of DDSPNs are presented comprising the automatic detection of conflicts and confusions, which removes a major obstacle for the analysis of discrete time models.
Approximations to Bayes estimate for quantal assay with simple exponential tolerance distribution
A discrete and, as approximation to it, a continuous model for the software reliability growth process are examined. The discrete model is based on independent multinomial trials and concerns itself with the joint distribution of the first occurrence time of its underlying events (bugs). The continuous model is based on the order statistics of N independent nonidentically distributed exponential random variables. It is shown that the spacings between bugs are not necessarily independent or exponentially (geometrically) distributed. However, there is a statistical rationale for viewing them so conditionally. Some identifiability problems are pointed out and resolved. In particular, it appears that the number of bugs in a program is not identifiable. Estimated upper bounds and confidence bounds for the residual program eror content are given based on the spacings of the first k bugs removed.
Dislocations are line defects in crystals that multiply and self-organize into a complex network during strain hardening. The length of dislocation links, connecting neighboring nodes within this network, contains crucial information about the evolving dislocation microstructure. By analyzing data from Discrete Dislocation Dynamics (DDD) simulations in face-centered cubic (fcc) Cu, we characterize the statistical distribution of link lengths of dislocation networks during strain hardening on individual slip systems. Here, our analysis reveals that link lengths on active slip systems follow a double-exponential distribution, while those on inactive slip systems conform to a single-exponential distribution. The distinctive long tail observed in the double-exponential distribution is attributed to the stress-induced bowing out of long links on active slip systems, a feature that disappears upon removal of the applied stress. We further demonstrate that both observed link length distributions can be explained by extending a one-dimensional Poisson process to include different growth functions. Specifically, the double-exponential distribution emerges when the growth rate for links exceeding a critical length becomes super-linear, which aligns with the physical phenomenon of long links bowing out under stress. This work advances our understanding of dislocation microstructure evolution during strain hardening and elucidates the underlying physical mechanisms governing its formation.
Abstract Robust principal component analysis (RPCA) is a widely used method for recovering low‐rank structure from data matrices corrupted by significant and sparse outliers. These corruptions may arise from occlusions, malicious tampering, or other causes for anomalies, and the joint identification of such corruptions with low‐rank background is critical for process monitoring and diagnosis. However, existing RPCA methods and their extensions largely do not account for the underlying probabilistic distribution for the data matrices, which in many applications are known and can be highly non‐Gaussian. We thus propose a new method called RPCA for exponential family distributions (), which can perform the desired decomposition into low‐rank and sparse matrices when such a distribution falls within the exponential family. We present a novel alternating direction method of multiplier optimization algorithm for efficient decomposition, under either its natural or canonical parametrization. The effectiveness of is then demonstrated in two applications: the first for steel sheet defect detection and the second for crime activity monitoring in the Atlanta metropolitan area.
The ability to estimate the fraction of ground flashes in a set of flashes observed by a satellite lightning imager, such as the future GOES-R Geostationary Lightning Mapper (GLM), would likely improve operational and scientific applications (e.g., severe weather warnings, lightning nitrogen oxides studies, and global electric circuit analyses). A Bayesian inversion method, called the Ground Flash Fraction Retrieval Algorithm (GoFFRA), was recently developed for estimating the ground flash fraction. The method uses a constrained mixed exponential distribution model to describe a particular lightning optical measurement called the Maximum Group Area (MGA). To obtain the optimum model parameters (one of which is the desired ground flash fraction), a scalar function must be minimized. This minimization is difficult because of two problems: (1) Label Switching (LS), and (2) Parameter Identity Theft (PIT). The LS problem is well known in the literature on mixed exponential distributions, and the PIT problem was discovered in this study. Each problem occurs when one allows the numerical minimizer to freely roam through the parameter search space; this allows certain solution parameters to interchange roles which leads to fundamental ambiguities, and solution error. A major accomplishment of this study is that we have employed a state-of-the-art genetic-based global optimization algorithm called Differential Evolution (DE) that constrains the parameter search in such a way as to remove both the LS and PIT problems. To test the performance of the GoFFRA when DE is employed, we applied it to analyze simulated MGA datasets that we generated from known mixed exponential distributions. Moreover, we evaluated the GoFFRA/DE method by applying it to analyze actual MGAs derived from low-Earth orbiting lightning imaging sensor data; the actual MGA data were classified as either ground or cloud flash MGAs using National Lightning Detection Network[TM] (NLDN) data. Solution error plots are provided for both the simulations and actual data analyses.
We show that, although nonlinear optics may give rise to a vast multitude of statistics, all these statistics converge, in their extreme-value limit, to one of a few universal extreme-value statistics. Specifically, in the class of polynomial nonlinearities, such as those found in the Kerr effect, weak-field harmonic generation, and multiphoton ionization, the statistics of the nonlinear-optical output converges, in the extreme-value limit, to the exponentially tailed, Gumbel distribution. Exponentially growing nonlinear signals, on the other hand, such as those induced by parametric instabilities and stimulated scattering, are shown to reach their extreme-value limits in the class of the Fréchet statistics, giving rise to extreme-value distributions (EVDs) with heavy, manifestly nonexponential tails, thus favoring extreme-event outcomes and rogue-wave buildup.
The relative abundances are treated as a consequence of processes in cosmic ray transport occurring during passage of the radiation through interstellar material at high velocity. Some of the subjects mentioned are nuclear fragmentation and the production of secondary nuclei, nuclear reactions, energy loss and nuclear decay, ionization, the range-energy relation and propagation variables, capture and loss of electrons, the propagation of nuclei, the transport equation, equilibrium solutions, energy-dependent path length distribution, exponential path length distributions, discrete spectra, sources, supernovae, and the origin of the abundances. The connection between the space-time features of the sources, the material traversed, and the effects of magnetic fields is established by describing the particle-field interaction as a diffusive or random-walk process.
In the bus network problem, the goal is to generate a plan for getting from point X to point Y within a city using buses in the smallest expected time. Because bus arrival times are not determined by a fixed schedule but instead may be random. the problem requires more than standard shortest path techniques. In recent work, Datar and Ranade provide algorithms in the case where bus arrivals are assumed to be independent and exponentially distributed. We offer solutions to two important generalizations of the problem, answering open questions posed by Datar and Ranade. First, we provide a polynomial time algorithm for a much wider class of arrival distributions, namely those with increasing failure rate. This class includes not only exponential distributions but also uniform, normal, and gamma distributions. Second, in the case where bus arrival times are independent and geometric discrete random variable,. we provide an algorithm for transportation networks of buses and trains, where trains run according to a fixed schedule.
The high resolution and global coverage of the Magellan radar image data set allows detailed study of the smallest volcanoes on the planet. A modified classification scheme for volcanoes less than 20 km in diameter is shown and described. It is based on observations of all members of the 556 significant clusters or fields of small volcanoes located and described by this author during data collection for the Magellan Volcanic and Magmatic Feature Catalog. This global study of approximately 10 exp 4 volcanoes provides new information for refining small volcano classification based on individual characteristics. Total number of these volcanoes was estimated to be 10 exp 5 to 10 exp 6 planetwide based on pre-Magellan analysis of Venera 15/16, and during preparation of the global catalog, small volcanoes were identified individually or in clusters in every C1-MIDR mosaic of the Magellan data set. Basal diameter (based on 1000 measured edifices) generally ranges from 2 to 12 km with a mode of 34 km, and follows an exponential distribution similar to the size frequency distribution of seamounts as measured from GLORIA sonar images. This is a typical distribution for most size-limited natural phenomena unlike impact craters which follow a power law distribution and continue to infinitely increase in number with decreasing size. Using an exponential distribution calculated from measured small volcanoes selected globally at random, we can calculate total number possible given a minimum size. The paucity of edifice diameters less than 2 km may be due to inability to identify very small volcanic edifices in this data set; however, summit pits are recognizable at smaller diameters, and 2 km may represent a significant minimum diameter related to style of volcanic eruption. Guest, et al, discussed four general types of small volcanic edifices on Venus: (1) small lava shields; (2) small volcanic cones; (3) small volcanic domes; and (4) scalloped margin domes ('ticks'). Steep-sided domes or 'pancake domes', larger than 20 km in diameter, were included with the small volcanic domes. For the purposes of this study, only volcanic edifices less than 20 km in diameter are discussed. This forms a convenient cutoff since most of the steep-sided domes ('pancake domes') and scalloped margin domes ('ticks') are 20 to 100 km in diameter, are much less numerous globally than are the smaller diameter volcanic edifices (2 to 3 orders of magnitude lower in total global number), and do not commonly occur in large clusters or fields of large numbers of edifices.
Frequency histograms and the 'power spectrum analysis' (PSA) method, the latter developed by Yu & Peebles (1969), have been widely employed as techniques for establishing the existence of periodicities. We provide a formal analysis of these two classes of methods, including controlled numerical experiments, to better understand their proper use and application. In particular, we note that typical published applications of frequency histograms commonly employ far greater numbers of class intervals or bins than is advisable by statistical theory sometimes giving rise to the appearance of spurious patterns. The PSA method generates a sequence of random numbers from observational data which, it is claimed, is exponentially distributed with unit mean and variance, essentially independent of the distribution of the original data. We show that the derived random processes is nonstationary and produces a small but systematic bias in the usual estimate of the mean and variance. Although the derived variable may be reasonably described by an exponential distribution, the tail of the distribution is far removed from that of an exponential, thereby rendering statistical inference and confidence testing based on the tail of the distribution completely unreliable. Finally, we examine a number of astronomical examples wherein these methods have been used giving rise to widespread acceptance of statistically unconfirmed conclusions.
In the surface hydrologic parameterization of general circulation models (GCMs), it is commonly assumed that the precipitation processes are homogeneous over a GCM grid square and that the precipitation intensity is uniformly distributed. Based on evidence that the spatial distribution of precipitation within a GCM grid square is crucial for the land surface hydrology parameterization, a few researchers have explored the impacts of assuming that the precipitation is exponentially distributed. This paper explores the suitability of the aforementioned assumptions. First, a statistical analysis is conducted of historical precipitation data for three GCM grids in different regions of the United States. The analysis suggests that neither the uniform nor the exponential distribution assumption may be suitable at the GCM grid scale and, that instead, the spatial variability in precipitation is characterized by statistical patterns that are inhomogeneous. These patterns vary from grid to grid and are induced by the interaction between atmospheric conditions and various land surface characteristics, such as topographical features, surface properties, etc. Within the same grid square, however, the statistical patterns are generally constant from year to year. Based on this analysis, a computationally viable (i.e., usable with GCMs) stochastic precipitation disaggregation scheme that utilizes these stable statistical patterns is proposed. The method was used to generate spatially distributed hourly rainfall for a summer season in the southwestern region of the continental United States. Analysis of the results shows that the methodology preserves the seasonal characteristics of spatial variability in precipitation that is observed in the long-term historical data.
During the early development of quantum chromodynamics, it was proposed that baryon number could be carried by a non-perturbative Y-shaped topology of gluon fields, called the gluon junction, rather than by the valence quarks as in the QCD standard model. A puzzling feature of ultra-relativistic nucleus-nucleus collisions is the apparent substantial baryon excess in the mid-rapidity region that could not be adequately accounted for in most conventional models of quark and diquark transport. The transport of baryonic gluon junctions is predicted to lead to a characteristic exponential distribution of net-baryon density with rapidity and could resolve the puzzle. In this context we point out that the rapidity density of net-baryons near mid-rapidity indeed follows an exponential distribution with a slope of –0.61 ± 0.03 as a function of beam rapidity in the existing global data from A+A collisions at AGS, SPS and RHIC energies. To further test if quarks or gluon junctions carry the baryon quantum number, we propose to study the absolute magnitude of the baryon vs. charge stopping in isobar collisions at RHIC. We also argue that semi-inclusive photon-induced processes (γ + p/A) at RHIC kinematics provide an opportunity to search for the signatures of the baryon junction and to shed light onto the mechanisms of observed baryon excess in the mid-rapidity region in ultra-relativistic nucleus-nucleus collisions. Such measurements can be further validated in A+A collisions at the LHC and e + p/A collisions at the EIC.
Abstract Perpendicular magnetic tunnel junction (pMTJ)-based true-random number generators (RNGs) can consume orders of magnitude less energy per bit than CMOS pseudo-RNGs. Here, we numerically investigate with a macrospin Landau–Lifshitz-Gilbert equation solver the use of pMTJs driven by spin–orbit torque to directly sample numbers from arbitrary probability distributions with the help of a tunable probability tree. The tree operates by dynamically biasing sequences of pMTJ relaxation events, called ‘coinflips’, via an additional applied spin-transfer-torque current. Specifically, using a single, ideal pMTJ device we successfully draw integer samples on the interval [0, 255] from an exponential distribution based on p -value distribution analysis. In order to investigate device-to-device variations, the thermal stability of the pMTJs are varied based on manufactured device data. It is found that while repeatedly using a varied device inhibits ability to recover the probability distribution, the device variations average out when considering the entire set of devices as a ‘bucket’ to agnostically draw random numbers from. Further, it is noted that the device variations most significantly impact the highest level of the probability tree, with diminishing errors at lower levels. The devices are then used to draw both uniformly and exponentially distributed numbers for the Monte Carlo computation of a problem from particle transport, showing excellent data fit with the analytical solution. Finally, the devices are benchmarked against CMOS and memristor RNGs, showing faster bit generation and significantly lower energy use.