Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Bayes methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Stresses in a two-bay noncircular cylinder under transverse loads

A method, taking into account the effects of flexibility and based on a general eighth-order differential equation, is presented for finding the stresses in a two-bay, noncircular cylinder the cross section of which can be composed of circular arcs. Numerical examples are given for two cases of ring flexibility for a cylinder of doubly symmetrical (essentially elliptic) cross section, subjected to concentrated radial, moment, and tangential loads. The results paralleled those already obtained for shells with circular rings.

Griffith, George E↗

Front Delineation and Tracking with Multiple Underwater Vehicles

This work describes a method for detecting and tracking ocean fronts using multiple autonomous underwater vehicles. Multiple vehicles — equally-spaced along the expected frontal boundary — complete near parallel transects orthogonal to the front. Lateral gradients are used to determine the location of the front crossing from each individual vehicle transect by detecting a change in the observed water property. Adaptive control of the vehicles ensure they remain perpendicular to the estimated front boundary as it evolves over time. This method was demonstrated in and around Monterey Bay, California in May of 2017. We compare the front detection method to previously used methods. We introduce a metric in order to evaluate the adaptive control techniques presented. We show the capability of this method for repeated sampling across a dynamic two-dimensional ocean front using short-range Iver AUVs. This method extends to tracking gradients of different properties using a variety of vehicles.

Chavez, Francisco P.↗

Structural damage detection of space truss structures using best achievable eigenvectors

A method is presented by which measured modes and frequencies from a modal test can be used to determine the location and magnitude of damage in a space struss structure. The damage is located by computing the Euclidean distances between the measured mode shapes and the best achievable eigenvectors. The best achievable eigenvectors are the projection of the measured mode shapes onto the subspace defined by the refined analytical model of the structure and the measured frequencies. Loss of both stiffness and mass properties can be located and quantified. To examine the performance of the method when experimentally measured modes are employed, various damage detection studies using a laboratory eight-bay truss structure were conducted. The method performs well even though the measurement errors inevitably make the damage location more difficult.

Lim, Tae W.↗

Cost Efficiency of Environmental DNA as Compared to Conventional Methods for Biodiversity Monitoring Purposes at Marine Energy Sites

The installation of marine energy systems may affect marine environments, and by extension, marine fish communities. Therefore, biomonitoring is an integral part of assessing impacts on species. Environmental DNA (eDNA) provides a noninvasive alternative to conventional monitoring surveys and the possibility of a more accurate assessment of species richness. Yet, its cost efficiency compared to traditional methods of monitoring is relatively unknown, especially when applied to monitoring around tidal, wave, and offshore wind energy installations. For this study, 202 peer-reviewed journal articles were dissected to inventory the diversity of supplies used for collecting and processing eDNA samples and to compile the average cost of eDNA surveys. Information collected included the type, volume, and brand of containers used in sampling; material, size, and brand of filters; and extraction methods. Cost information was gathered for the most common supplies, and a total cost was estimated for a hypothetical eDNA survey in Sequim Bay, WA, to compare with traditional methods of surveying such as beach seining and scuba surveys. The results showed a higher-than-expected diversity of supplies to collect and process eDNA samples. The most common supplies were 1 L Nalgene bottles at an average cost of 7.96 USD for collecting samples, 0.45 µm glass fiber Merck Millipore filters at an average cost of 1.51 USD for filtering samples, and the Qiagen DNeasy Blood and Tissue kit at 3.54 USD per sample for extracting DNA. When compared to beach seine and scuba surveys, eDNA surveys undertaken by senior researchers are less expensive for both initial surveys with all new materials as well as for follow-up surveys reusing some of the supplies. However, when surveys are done solely by students, eDNA surveys are more expensive than scuba surveys when no prior supplies are available and more than both beach seine and scuba surveys for follow-up surveys reusing supplies. In a professional sphere, where surveys are less often conducted by teams of students only, eDNA surveys are an effective and less-costly alternative to conventional methods. We anticipate that the development and refinement of eDNA methodology will continue to decrease surveying costs.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Command and data handling of science signals on Spacelab

The Orbiter Avionics and the Spacelab Command and Data Management System (CDMS) combine to provide a relatively complete command, control, and data handling service to the instrument complement during a Shuttle Sortie Mission. The Spacelab CDMS services the instruments and the Orbiter in turn services the Spacelab. The CDMS computer system includes three computers, two I/O units, a mass memory, and a variable number of remote acquisition units. Attention is given to the CDMS high rate multiplexer, CDMS tape recorders, closed circuit television for the visual monitoring of payload bay and cabin area activities, methods of science data acquisition, questions of transmission and recording, CDMS experiment computer usage, and experiment electronics.

Mccain, H. G.↗

Bayesian estimation of crack initiation times from service data

Lockheed C-130 Hercules aircraft have during their service life been periodically inspected and growing cracks around rivet holes were recorded. This record has recently been used to determine the statistical distributions of crack initiation times and the distribution of initial crack sizes. When crack initiation times are calculated from such cracks, by backward extrapolation of the growth relation, the resulting distribution of crack initiation times will indicate a preponderance of short times to crack initiation. If however, such distributions are combined with the reliability of the inspection procedure, the statistical distribution of missed initiation times can be estimated. The method used is based on Bayes theorem which permits the calculation of the 'prior' distribution (initiation times before inspection) from a knowledge of the 'posterior' distribution (initiation times obtained from the inspection) and a 'likelihood function' (reliability of the inspection) procedure. The results indicate that during an early inspection a large percentage of initiation times will be missed and that the fraction of located initiation times increases during later inspections.

Heller, R. A.↗

The Chesapeake Bay Program: An opportunity to use an innovative monitoring technique

The goal of this program is to develop a management system that will protect and preserve the water quality of the Chesapeake Bay by effectively managing its uses and resources. To achieve this goal, three major objectives must be accomplished: (1) Determine what units of government have management responsibility for the environmental quality of the Chesapeake Bay, also to define how such management responsibility can best be structured so that communications and coordination can be improved between the respective units of government, research, educational institutions, concerned groups, and individuals. (2) Assess the principal factors having an adverse impact on the environmental quality of the Chesapeake Bay. Following this assessment and review of ongoing research, direct and coordinate research and abatement programs that will most effectively address these factors, and (3) analyze all environmental sampling data now being collected on the Chesapeake Bay and suggest and undertake methods for improving this data collection, and to establish a continuing capability for collecting storing, analyzing, and disseminating these data.

Mangiaracina, L.↗

Lightning observations from space shuttle

The experimental program of the Earth Sciences and Applications Division at NASA/MSFC includes development of the Lightning Imaging Sensor (LIS) for the NOAA Earth Observing System (EOS) Polar Platform. The research plan is to use existing lightning information to generate simulated data for the LIS experiment. Navigation algorithms were used to transform pixel locations to latitude and longitude values. The simulated data would then be used to test and develop algorithms for the analysis of LIS data. Individual frames of video imagery obtained from Space Shuttle Missions provide the raw data for the simulation. Individual video frames were digitized to get the pixel locations of lightning flashes. The pixel locations will be used to locate the geographical position of the event. Because of a lack of detailed knowledge of camera orientation with respect to the Space Shuttle, video scenes that contain identifiable city lights were chosen for analysis. A method for locating the payload bay camera axis was developed and tested. Two measurements are needed: the pixel location of the apparent horizon and a timed siting of a known location passing the principal line of the image. Individual video frames were navigated and lightning illuminated clouds were located on the map. Satisfactory agreement in location was achieved for cities and LLP lightning locations. Ground truth measurements were compared to satellite observations. A vertical lightning event was identified on the horizon. Very low frequency (VLF) transmission on this particular occassion shows a strong response to negative cloud to cloud flashes.

Boeck, William L.↗

A Machine Learning Approach to Improve Air Traffic Management Initiatives

Collaborating closely with commercial air carriers and related organizations, the Federal Aviation Administration(FAA) regulates air traffic and ensures the safety and efficiency of air operations. Air traffic controllers make strategic decisions, such as delaying, rerouting, or canceling flights, partly based on guidance provided by the FAA’s Air TrafficControl System Command Center (ATCSCC). The guidance includes, among other things, control measures known asTraffic Management Initiatives (TMIs) designed to enhance safety and improve operational efficiency. TMIs play a crucial role in managing the demand and capacity within the U.S. National Airspace System (NAS). Two major TMIs that are routinely used (primarily to mitigate the adverse effects of bad weather) are Ground Delay Programs (GDPs) andGround Stops (GSs). In a GDP, flights destined for airports facing thunderstorm activity experience delays at their origin airports. This proactive approach minimizes the risk of routing aircraft through hazardous weather conditions and also replaces (fuel burning) airborne delays with ground delays. In a GS, a temporary restriction is imposed on the departure or arrival of aircraft at a specific airport or within a designated airspace. Although other TMIs (e.g., miles-in-trail) are also implemented as part of (air) traffic flow management in the NAS, the focus of this work is on GDPs and GSs. Since TMIs, by design, lead to flight delays or cancellations, it is crucial to put in place the right set of parameters(e.g., scope and duration of the GDP). For example, when the end time of a GDP extends beyond what is necessary, it imposes unnecessary delays on departing flights. This situation could occur as a result of inaccurate prediction of the(required) duration of the GDP based on the weather forecast. On the other hand, if a GDP ends prematurely before the underlying capacity constraints are resolved at the destination airport, it may result in airborne holding. The delicate balance lies in matching the termination of the GDP precisely with the resolution of capacity constraints, avoiding both the imposition of unnecessary ground delays and the need for airborne holding due to premature program termination.Failing to specify the right parameters for TMIs also leads to flight delays, creating a significant obstacle in managing the increasing traffic volumes causing increased work load for the controllers. To address this issue, we propose the integration of Machine Learning (ML) models in the traffic flow management(TFM) pipeline. In current operations, decisions are made by human experts based on extensive training, historical patterns, available traffic and weather data. Since we have an abundance of data from past events that tell us the likely impact of various TMIs, by ingesting historical data, properly trained ML models can offer valuable insights and aid human decision-making. With the FAA increasingly exploring advanced analytics, ML emerges as a focal point for enhancing TFM within the National Airspace System (NAS). As a first step, this study aims to provide traffic controllers with decision-making support for the issuance and adjustment of TMIs. Data analytics and machine learning have been previously employed to address some of the challenges associated with TMIs. Numerous studies have concentrated on various facets of TMI issuance, exploring factors influencing TMI parameters, including arrival rate, airport capacity, and delay prediction. For example, using weather forecasts, several statistical methods were used to produce probabilistic capacity profiles which in conjunction with deterministic models provided insights into the GDP planning process [1–4]. The downside of using deterministic models is that they rely on fixed inputs and predetermined rules, which lack the ability to account for the inherent uncertainty and variability present in real-world scenarios. In a separate series of studies, researchers aimed to predict the occurrences of GDPs and GSs. The majority of these studies utilized various supervised learning methods, including Decision Trees, Naive Bayes, Support VectorMachines, and Random Forests to analyze the influence of weather conditions and arrival demand on TMI incidents[5–8]. However, these studies primarily focused on predicting the incidence of TMIs without explicitly addressing the scope of TMIs, including their duration and their geographical coverage. Furthermore, the emphasis of these studies was largely on GDPs, given their higher frequency and longer duration when compared to GSs. A limited number of studies focused on predicting the parameters of TMIs, specifically addressing their duration and extent. In one such study focusing on optimizing the TMI parameters at San Francisco International Airport (SFO),the authors utilized a probabilistic forecast of fog [9]. They simulated various capacity scenarios based on the (fog)burn-off forecasts, selecting GDP parameters that minimized airborne and overall ground delays. However, this approach exclusively emphasizes stratus (fog) burn-off as the primary determinant of GDP and GS, neglecting other influential factors like severe weather events, runway closures, lower capacity than traffic demand, and other important variables. Given the complexity of predicting the TMI and determining its scope, we seek a more holistic approach. We aim to consider all significant factors that could impact TMIs and their parameters. What sets this research apart is the fusion of all data sources relevant to the issuance and adjustment of TMIs and it represents the first comprehensive attempt to optimize TMIs in this manner. Since this comprehensive solution involves various aspects, we break down the problem into smaller components and input all parameters into a unified model called the “TMI Adjuster”. Figure 1 shows the overall framework and the list of datasets used in each model. The objective of the TMI Adjuster module is to deliver reliable, consistent and expedited recommendations for the progression, adjustment, and termination of TMIs. The ML solution entails developing a pipeline capable of predicting the necessity of a TMI (e.g., GS or GDP) along with its various parameters. For example, in the case of a GS, this includes the scope of the GS either in terms of distance from the destination airport or based on pre-defined airspace sectors. Here, scope refers to those regions and departing airports that are subject to the GS. In this paper, we concentrate on the issuance of GSs in the three major airports in the New York area — LaGuardia(LGA), John F. Kennedy International (JFK), and Newark Liberty International (EWR). We fuse traffic, weather and other relevant aviation data from years 2017 to 2019 to train and validate the ML models. In particular, we use the following datasets: •Terminal Aerodrome Forecast (TAF): meteorological forecasts specific to each airport, issued four times a day, covering predefined time periods. •TMI data: includes all GSs and GDPs along with their respective parameters. •Aviation System Performance Metrics (ASPM): includes traffic related data such as aircraft delays, arrival, and departure rates. •Notices to Airmen (NOTAMs): utilized to extract runway closure data and manage interdependencies between terminals in close proximity. •Flight cancellation data •Airspace Flow Programs (AFP): includes information on flight airborne holdings caused by TMIs. The data preprocessing entails transforming ASPM, TMI, AFP, NOTAMs, and weather data into an hourly format and consolidating all datasets by merging them based on date and time as the primary key. The TMI Adjuster framework comprises two parallel models: one dedicated to GS and a second model focused on GDP. As previously mentioned, our specific focus is on the GS model as a multi-classification problem. In this framework, each data point of the GS model input summarizes ten hours of data. Specifically, the data loader for the GS model generates the input and output of the model as follows: at a given time step, the input includes the actual traffic, weather, and TMI data from the two-hour window before the time step, alongside the weather forecast and scheduled traffic for the next 8 hours starting from the time step. Based on this information, the output of the GS model for each time interval consists of three dimensions. The first dimension represents a binary decision on whether there should be a GS in place for the next hour or not. The second dimension is related to the scope of the GS in the United States, and the third dimension is related to the scope of the GS in Canada (i.e., to determine if the GS impacts airports in Canada).One of the challenges with TMI modeling is the sparsity of TMI events, particularly regarding its scope. To address this challenge in the scope of the GS model output, we implement grouping. The GS scope for the US region is defined based on a list of centers that should be included when the GS is in place. With 20 centers in the US, we utilized historical data to group them into 4 categories. In particular, we summarized our historical data in a graph format where nodes represent centers, and link weights are defined based on the co-occurrence of centers in the scope parameter ofTMIs. By identified strongly connected components in this graph, we were able to partition the centers into four groups. We consider two model structures for the GS Model. Firstly, a hierarchical classification model [10], where the human decision-making for a GS is of hierarchical nature. The decision-maker first decides whether there is a need fora GS, and if the answer is yes, determines the scope. A hierarchical classification model organizes the problem into a class hierarchy, typically a tree or a Directed Acyclic Graph (DAG) structure, and considers the dependency of the decision in the previous step to the next component [10]. Here, we employ the local classifier per level approach, which involves training one multi-class classifier for each level of the class hierarchy. The second structure is the independent structure. In this setting, as the name suggests, we do not consider the dependency of the decisions in the different dimensions of the output of the model. Instead, for each dimension, we train a multi-class classifier independently. Table 1 summarizes GS model statistics for training, validation and testing. The table documents the effect of limiting data to the time steps when there was actually a TMI in place or when a TMI had just terminated. This resulted in a more balanced distribution of the GS class(GS positive class)versus “No GS”(GS negative class), which might help the training process. While JFK and LGA follow very similar distributions, with 40% and 42% GS positive class respectively, EWR has proportionally fewer GS incidents at 28%. Our subsequent phase involves evaluating the performance of both hierarchical structure and independent structure using different state-of-the-art multi-class classifier models such as Random Forest, Decision Trees, K-nearest Neighbors, and Logistic Regression and forecast the duration and scope of the GSs.

Farzan Masrour Shalmani↗

Robustness analysis applied to substructure controller synthesis

The stability and robustness of the controlled system obtained via the substructure control synthesis (SCS) method of Su et al. (1990) were examined using a six-bay truss model, and employing an LQG control design method to obtain controllers for two separate structures. It is found that the assembled controller provides a stability in this instance. A qualitative assessment of the stability robustness of the system with controller designed with the SCS method is provided by obtaining a controller using the complete truss model and comparing the robustness of the corresponding closed-loop systems.

Gonzalez-Oberdoerffer, Marcelo F.↗

Determination of proton PDF uncertainties with Markov chain Monte Carlo

We present an analysis of parton distribution functions (PDFs) of the proton using Markov chain Monte Carlo (MCMC) methods. The MCMC approach naturally implements Bayes’ theorem and, thus, provides a means to directly sample the underlying probability distribution—in this case, the probability distribution of the PDF parameters. This allows for a straightforward propagation of the resulting uncertainties into any PDF-dependent observable, preserving their simple probabilistic interpretation. In our analysis we include a broad set of deep inelastic scattering data from HERA, BCDMS and NMC experiments along with the Drell-Yan, 𝑊 and 𝑍 boson data from LHC and Tevatron experiments, which combined with theoretical calculations at next-to-next-to-leading order in QCD allow for realistic determination of PDFs. The main focus of this analysis is to explore alternative methods for PDF uncertainty estimation that are more firmly grounded in statistical principles. We show that the flexibility of the Bayes framework, allowing one, e.g., to account for non-Gaussianity or inconsistencies of datasets, is crucial to extract realistic uncertainties when such assumptions are not fulfilled. We also demonstrate that MCMC allows one to determine the Δ⁢𝜒 2 value corresponding to a given confidence level in the sample, which can, in turn, be used as a statistically well-founded tolerance criterion used in the Hessian method, thus addressing one of its main long-standing drawbacks.

Risse, Peter Clemens [Universität Münster (Germany↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), has typically limited machine learning (ML) in space studies and further study of radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNAseq) data from 6 mouse liver GeneLab datasets (GLDS) with a total of 113 spaceflight and ground-control samples to determine top features relevant to spaceflight including the effect of radiation exposure. Data was normalized within each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. The top MRMR features were used to predict spaceflight vs. ground-control samples using a Random Forest (RF) classifier with 5-fold cross validation (CV). The ML-based gene sets were further compared against differential gene expression results from individual GLDS. CV training using the top 100 MRMR genes show averages of 86% accuracy and 0.95 AUC value on the validation set over 5 folds (Figure 1A). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 811 or 68 DEGs overlapping between at least 2 or 3 studies, respectively (Figure 1B). Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism. Set analysis between the MRMR features and the DEGs showed 60 or 8 genes overlapping with at least 1 or 2 studies, respectively. MRMR feature selection and ensemble ML methods (e.g. RF) improve performance relative to a Naïve Bayes classifier when NGS data sets are analyzed. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise ratio. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from RNASeq analysis. Non-intersecting sets introduce opportunity to explore spaceflight relevant genes and implementing ML methods across existing NGS datasets may overcome sample size limitations. ML coupled with existing analytical methods enhances understanding of disease by revealing common underlying pathways across datasets.

Machine Learning↗

Water resources planning for rivers draining into Mobile Bay

The application of remote sensing, automatic data processing, modeling and other aerospace related technologies to hydrological engineering and water resource management are discussed for the entire river drainage system which feeds the Mobile Bay estuary. The adaptation and implementation of existing mathematical modeling methods are investigated for the purpose of describing the behavior of Mobile Bay. Of particular importance are the interactions that system variables such as river flow rate, wind direction and speed, and tidal state have on the water movement and quality within the bay system.

April, G. C.↗

Permanent Sequestration of Emitted Gases in the Form of Clathrate Hydrates

Underground sequestration has been proposed as a novel method of permanent disposal of harmful gases emitted into the atmosphere as a result of human activity. The method was conceived primarily for disposal of carbon dioxide (CO2, greenhouse gas causing global warming), but could also be applied to CO, H2S, NOx, and chorofluorocarbons (CFCs, which are super greenhouse gases). The method is based on the fact that clathrate hydrates (e.g., CO2 6H2O) form naturally from the substances in question (e.g., CO2) and liquid water in the pores of sub-permafrost rocks at stabilizing pressures and temperatures. The proposed method would be volumetrically efficient: In the case of CO2, each volume of hydrate can contain as much as 184 volumes of gas. Temperature and pressure conditions that favor the formation of stable clathrate hydrates exist in depleted oil reservoirs that lie under permafrost. For example, CO2-6H2O forms naturally at a temperature of 0 C and pressure of 1.22 MPa. Using this measurement, it has been calculated that the minimum thickness of continuous permafrost needed to stabilize CO2 clathrate hydrate is only about 100 m, and the base of the permafrost is known to be considerably deeper at certain locations (e.g., about 600 m at Prudhoe Bay in Alaska). In this disposal method, the permafrost layers over the reservoirs would act as impermeable lids that would prevent dissociation of the clathrates and diffusion of the evolved gases up through pores.

Duxbury, N.↗

Analysis of Waves in Space Plasma (WISP) near field simulation and experiment

The WISP payload scheduler for a 1995 space transportation system (shuttle flight) will include a large power transmitter on board at a wide range of frequencies. The levels of electromagnetic interference/electromagnetic compatibility (EMI/EMC) must be addressed to insure the safety of the shuttle crew. This report is concerned with the simulation and experimental verification of EMI/EMC for the WISP payload in the shuttle cargo bay. The simulations have been carried out using the method of moments for both thin wires and patches to stimulate closed solids. Data obtained from simulation is compared with experimental results. An investigation of the accuracy of the modeling approach is also included. The report begins with a description of the WISP experiment. A description of the model used to simulate the cargo bay follows. The results of the simulation are compared to experimental data on the input impedance of the WISP antenna with the cargo bay present. A discussion of the methods used to verify the accuracy of the model is shown to illustrate appropriate methods for obtaining this information. Finally, suggestions for future work are provided.

Richie, James E.↗

Role of remote sensing in Bay measurements

Remote measurements of a number of surface or near surface parameters for baseline definition and specialized studies, remote measurements of episodic events, and remote measurements of the Bay lithosphere are considered in terms of characterizing and understanding the ecology of the Chesapeake Bay. Geologic processes and features best suited for information enhancement by remote sensing methods are identified. These include: (1) rates of sedimentation in the Bay; (2) rates of erosion of Bay shorelines; (3) spatial distribution and geometry of aquifers; (4) mapping of Karst terrain (sinkholes); and (5) mapping of fracture patterns. Recommendations for studying problem areas identified are given.

Mugler, J. P., Jr.↗

Textural analysis by statistical parameters and its application to the mapping of flow-structures in wetlands

From 1974 to 1977 the application of remote sensing methods in coastal areas and tidal bays and estuaries was investigated on the German coast of the North Sea. Aerial photographs were taken using different films; (1) color, (2) color infrared, and (3) black and white films. Scanner recordings were taken by an 11 channel scanner. Ground truth measurements of radiation and measurements of meteorological elements were carried out. For mapping the morphology in mudflat areas a digital texture analysis was developed, by which measurement of the change of image structures cased by distributing factors, such as changing illumination, is possible.

Wieczorek, U.↗

Bayesian Estimation of Earth’s Undiscovered Mineralogical Diversity Using Noninformative Priors

Recently, statistical distributions have been explored to provide estimates of the mineralogical diversity of Earth, and Earth-like planets. In this paper, a Bayesian approach is introduced to estimate Earth’s undiscovered mineralogical diversity. Samples are generated from a posterior distribution of the model parameters using Markov chain Monte Carlo simulations such that estimates and inference are directly obtained. It was previously shown that the mineral species frequency distribution conforms to a generalized inverse Gauss–Poisson (GIGP) large number of rare events model. Even though the model fit was good, the population size estimate obtained by using this model was found to be unreasonably low by mineralogists. In this paper, several zero-truncated, mixed Poisson distributions are fitted and compared, where the Poisson-lognormal distribution is found to provide the best fit. Subsequently, the population size estimates obtained by Bayesian methods are compared to the empirical Bayes estimates. Species accumulation curves are constructed and employed to estimate the population size as a function of sampling size. Finally, the relative abundances, and hence the occurrence probabilities of species in a random sample, are calculated numerically for all mineral species in Earth’s crust using the Poisson-lognormal distribution. These calculations are connected and compared to the calculations obtained in a previous paper using the GIGP model for which mineralogical criteria of an Earth-like planet were given.

Bayesian statistics↗