Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Multivariate Count”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

A Bayesian nonparametric analysis for zero-inflated multivariate count data with application to microbiome study

High-throughput sequencing technology has enabled researchers to profile microbial communities from a variety of environments, but analysis of multivariate taxon count data remains challenging. Here, we develop a Bayesian nonparametric (BNP) regression model with zero inflation to analyse multivariate count data from microbiome studies. A BNP approach flexibly models microbial associations with covariates, such as environmental factors and clinical characteristics. The model produces estimates for probability distributions which relate microbial diversity and differential abundance to covariates, and facilitates community comparisons beyond those provided by simple statistical tests. We compare the model to simpler models and popular alternatives in simulation studies, showing, in addition to these additional community-level insights, it yields superior parameter estimates and model fit in various settings. The model's utility is demonstrated by applying it to a chronic wound microbiome data set and a Human Microbiome Project data set, where it is used to compare microbial communities present in different environments.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Efficient GPU Implementation of Automatic Differentiation for Computational Fluid Dynamics

Many scientific and engineering applications require repeated calculations of derivatives of output functions with respect to input parameters. Automatic Differentiation (AD) is a method that automates derivative calculations and can significantly speed up code development. In Computational Fluid Dynamics (CFD), derivatives of flux functions with respect to state variables (Jacobian) are needed for efficient solutions of the nonlinear governing equations. AD of flux functions on graphics processing units (GPUs) is challenging as flux computations involve many intermediate variables that create high register pressure and require significant memory traffic because of the need to store the derivatives. This paper presents a forward-mode AD method based on multivariate dual numbers that addresses these challenges and simultaneously reduces the floating-point operation count. The dimension of the multivariate dual numbers is optimized for performance. The flux computations are restructured to minimize the number of temporary variables and reduce register pressure. For effective utilization of memory bandwidth, shared memory is used to store the local flux Jacobian. This AD implementation is compared with several other Jacobian implementations on an NVIDIA V100 GPU (V100). For three-dimensional perfect-gas compressible-flow equations implemented in a practical CFD code, the AD implementation of a flux Jacobian based on multivariate dual numbers of dimension 5 outperforms all other GPU AD implementations on V100. Its performance is comparable with the optimized hand-differentiated version. Finally, the implementation achieves 75% of the peak floating-point throughput and 61 % of the peak global device memory bandwidth usage.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Environmental exposure to industrial air pollution is associated with decreased male fertility

Objective: To understand how chronic exposure to industrial air pollution is associated with male fertility through semen parameters. Design: Retrospective cohort study. Subjects: Men in the Subfertility, Health and Assisted Reproduction cohort who underwent a semen analysis 2005-2017 with ≥1 measured semen parameter (N=21,563). Intervention(s): Residential histories for each man were constructed using locations from administrative records linked through the Utah Population Database. Industrial facilities with air emissions of nine endocrine disrupting compound chemical classes were identified from the Environmental Protection Agency Risk-Screening Environmental Indicators microdata. Chemical levels were linked with residential histories for the 5 years prior to each semen analysis. Main Outcome Measures: Semen analyses were classified as azoospermic or oligozoospermic (< 15 M/mL) using World Health Organization cutoffs for concentration. Bulk semen parameters such as concentration, total count, ejaculate volume, total motility, total motile count, and total progressive motile count were also measured. Multivariable regression models with robust standard errors were used to associate exposure quartiles for each of the nine chemical classes with each semen parameter, adjusting for age, race, and ethnicity, as well as neighborhood socioeconomic disadvantage. Results: After adjustment for demographic covariates, several chemical classes were associated with azoospermia and decreased total motility and volume. For exposure in the 4th relative to 1st quartile, significant associations were observed for acrylonitrile (β total motility = -0.87 pp), aromatic hydrocarbons (odds ratio [OR]azoospermia = 1.53; β volume = -0.14 mL), dioxins (OR azoospermia = 1.31; β volume = -0.09 mL; β total motility = -2.65 pp), heavy metals (β total motility = -2.78pp), organic solvents (OR azoospermia = 1.75; β volume = -0.10 mL), organochlorines (OR azoospermia = 2.09; β volume = -0.12 mL), phthalates (OR azoospermia = 1.44; β volume = -0.09 mL; β total motility = -1.21 pp), and silver particles (OR azoospermia = 1.64; β volume = -0.11 mL). All semen parameters significantly decreased with increasing socioeconomic disadvantage. Men who lived in the most disadvantaged areas had concentration, volume, and total motility of 6.70 M/mL, 0.13 mL, and 1.79 pp lower, respectively. Count, motile count, and total progressive motile count all decreased by 30–34 M.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

LandScan Global 30 Arcsecond Annual Global Gridded Population Datasets from 2000 to 2022

Abstract Oak Ridge National Laboratory (ORNL) annually develops the LandScan Global (LSG) dataset, a 30 arcsecond global gridded population dataset representing global ambient human population distribution. This multivariable dasymetric model disaggregates census counts within administrative boundaries using ancillary data. Each country’s distribution reflects cultural and socioeconomic patterns; manual validations yield a unique global dataset for assessing populations at risk. For over two decades, LSG has been a standard for estimating populations at risk, aiding U.S. federal government, academia and humanitarian organizations. During disasters such as the 2004 Indian Ocean tsunami and the 2010 Haiti earthquake and geopolitical crises such as the Syrian civil war and the 2022 Russian invasion of Ukraine, LSG supported scientific and operational communities in emergency response and recovery. In 2022, LSG datasets from 2000 onward were made publicly available through ORNL’s LandScan Portal. This data descriptor details our methodology and the application of geospatial science and machine learning to geographic and demographic data, highlighting uses in urban resiliency, emergency management, disaster response, and human health and security.

Science & Technology - Other Topics↗

A Step Beyond Simple Keyword Searches: Services Enabled by a Full Content Digital Journal Archive

The problems of managing and searching large archives of scientific journal articles can potentially be addressed through data mining and statistical techniques matured primarily for quantitative scientific data analysis. A journal paper could be represented by a multivariate descriptor, e.g., the occurrence counts of a number key technical terms or phrases (keywords), perhaps derived from a controlled vocabulary ( e . g . , the American Meteorological Society's Glossary of Meteorology) or bootstrapped from the journal archive itself. With this technique, conventional statistical classification tools can be leveraged to address challenges faced by both scientists and professional societies in knowledge management. For example, cluster analyses can be used to find bundles of "most-related" papers, and address the issue of journal bifurcation (when is a new journal necessary, and what topics should it encompass). Similarly, neural networks can be trained to predict the optimal journal (within a society's collection) in which a newly submitted paper should be published. Comparable techniques could enable very powerful end-user tools for journal searches, all premised on the view of a paper as a data point in a multidimensional descriptor space, e.g.: "find papers most similar to the one I am reading", "build a personalized subscription service, based on the content of the papers I am interested in, rather than preselected keywords", "find suitable reviewers, based on the content of their own published works", etc. Such services may represent the next "quantum leap" beyond the rudimentary search interfaces currently provided to end-users, as well as a compelling value-added component needed to bridge the print-to-digital-medium gap, and help stabilize professional societies' revenue stream during the print-to-digital transition.

Boccippio, Dennis J.↗

Crowd-based spatial risk assessment of urban flooding: Results from a municipal flood hotline in Detroit, MI

Climate change is increasing the frequency and intensity of extreme precipitation events, raising the risk of urban flood disasters. This study uses a crowd-sourced municipal call database to characterize the spatial distribution of flood risk in Detroit, MI. Call data including dates and addresses were obtained from the City of Detroit Department of Public Works for 2021. Calls were mapped and aggregated to census tract counts and merged with neighborhood-level data. Associations of predictors with flood calls were tested using spatial regression models. Flooding calls were located throughout the city but were concentrated in specific areas. Multivariate models of census tract level call counts indicated that increased poverty and Black, immigrant, and older residents were positively associated with flood calls, while increased elevation was associated with protective effects. Longer distances from waste water interceptors were associated with higher risk for calls. Crowd-sourced flood hotline call data can be used for effective spatial flood risk assessment. Though flooding occurs throughout the city of Detroit, infrastructural, neighborhood, and household factors influence flooding extent. Limitations included the self-reported nature of calls. Future modeling efforts might include input from local stakeholders to improve spatial risk assessment.

54 ENVIRONMENTAL SCIENCES↗

A Graphical Model for Fusing Diverse Microbiome Data

This paper develops a Bayesian graphical model for fusing disparate types of count data. The motivating application is the study of bacterial communities from diverse high-dimensional features, in this case, transcripts, collected from different treatments. In such datasets, there are no explicit correspondences between the communities and each corresponds to different factors, making data fusion challenging. We introduce a flexible multinomial-Gaussian generative model for jointly modeling such count data. This latent variable model jointly characterizes the observed data through a common multivariate Gaussian latent space that parameterizes the set of multinomial probabilities of the transcriptome counts. The covariance matrix of the latent variables induces a covariance matrix of co-dependencies between all the transcripts, effectively fusing multiple data sources. We present a computationally scalable variational Expectation-Maximization (EM) algorithm for inferring the latent variables and the parameters of the model. Here, the inferred latent variables provide a common dimensionality reduction for visualizing the data and the inferred parameters provide a predictive posterior distribution. In addition to simulation studies that demonstrate the variational EM procedure, we apply our model to a bacterial microbiome dataset.

59 BASIC BIOLOGICAL SCIENCES↗

Evaluating proxies for the drivers of natural gas productivity using machine-learning models

We report the extensive development of unconventional reservoirs using horizontal drilling and multistage hydraulic fracturing has generated large volumes of reservoir characterization and production data. The analysis of this abundant data using statistical methods and advanced machine-learning (ML) techniques can provide data-driven insights into well performance. Most predictive modeling studies have focused on the impact that different well completion and stimulation strategies have on well production but have not fully exploited the available in situ rock property data to determine its role in reservoir productivity. We have used machine-learning techniques to rank rock mechanical properties, microseismic attributes, and stimulation parameters in the order of their significance for predicting natural gas production from an unconventional reservoir. The data for this study came from a hydraulically fractured well in the Marcellus Shale in Monongalia County, West Virginia. The data classes included measurements aggregated by well completion stage that included (1) gas production, (2) well-log-derived measurements including bulk density, elastic moduli, shear impedance, compressional impedance, brittleness, and gamma measurements, (3) microseismic attributes, (4) long-period long-duration (LPLD) event counts, (5) fracture counts, and (6) stimulation parameters that included the fluid injection volume and average pumping pressure. To identify observable proxies for the drivers of gas production, we evaluated five commonly used ML approaches including multivariate adaptive regression spline, Gaussian mixture model, random forest, gradient boosting, and neural network. We selected five variables including LPLD event count, seismogenic b-value, hydraulic diffusivity, cumulative moment, and fluid volume as the features most likely to impact gas productivity at the stage level in the study area. The data-driven selection of these parameters for their importance in determining gas production can help reservoir engineers design more effective hydraulic-fracture treatments in the Marcellus Shale and other similar unconventional reservoirs. Plain language summary: We use machine-learning methods and data-driven selection of reservoir parameters to rank and better understand their importance in determining gas production, which can help reservoir engineers design more effective hydraulic-fracture treatments in the Marcellus Shale and other similar unconventional reservoirs.

58 GEOSCIENCES↗

The Poisson tensor completion parametric estimator

We introduce the Poisson tensor completion (PTC) estimator that exploits inter-sample relationships to compute a low-rank Poisson tensor decomposition of the frequency histogram for samples of a multivariate distribution. Our crucial observation is that the histogram bins are an instance of a space partitioning of counts and thus can be identified with a spatial non-homogeneous Poisson process. The Poisson tensor decomposition leads to a completion of the mean measure over all bins—including those containing few to no samples—and leads to our proposed estimator. A Poisson tensor decomposition models the underlying distribution of the count data and guarantees non-negative estimated values obviating the need for additional constraints to ensure non-negativity. Furthermore, we demonstrate that our PTC estimator is a substantial improvement over standard histogram-based estimators for sub-Gaussian probability distributions because of the concentration of norm phenomenon.

97 MATHEMATICS AND COMPUTING↗

Studies in Astronomical Time Series Analysis. VI. Bayesian Block Representations

This paper addresses the problem of detecting and characterizing local variability in time series and other forms of sequential data. The goal is to identify and characterize statistically significant variations, at the same time suppressing the inevitable corrupting observational errors. We present a simple nonparametric modeling technique and an algorithm implementing it-an improved and generalized version of Bayesian Blocks [Scargle 1998]-that finds the optimal segmentation of the data in the observation interval. The structure of the algorithm allows it to be used in either a real-time trigger mode, or a retrospective mode. Maximum likelihood or marginal posterior functions to measure model fitness are presented for events, binned counts, and measurements at arbitrary times with known error distributions. Problems addressed include those connected with data gaps, variable exposure, extension to piece- wise linear and piecewise exponential representations, multivariate time series data, analysis of variance, data on the circle, other data modes, and dispersed data. Simulations provide evidence that the detection efficiency for weak signals is close to a theoretical asymptotic limit derived by [Arias-Castro, Donoho and Huo 2003]. In the spirit of Reproducible Research [Donoho et al. (2008)] all of the code and data necessary to reproduce all of the figures in this paper are included as auxiliary material.

signal detection↗

Multivariate statistical analysis software technologies for astrophysical research involving large data bases

We developed a package to process and analyze the data from the digital version of the Second Palomar Sky Survey. This system, called SKICAT, incorporates the latest in machine learning and expert systems software technology, in order to classify the detected objects objectively and uniformly, and facilitate handling of the enormous data sets from digital sky surveys and other sources. The system provides a powerful, integrated environment for the manipulation and scientific investigation of catalogs from virtually any source. It serves three principal functions: image catalog construction, catalog management, and catalog analysis. Through use of the GID3* Decision Tree artificial induction software, SKICAT automates the process of classifying objects within CCD and digitized plate images. To exploit these catalogs, the system also provides tools to merge them into a large, complete database which may be easily queried and modified when new data or better methods of calibrating or classifying become available. The most innovative feature of SKICAT is the facility it provides to experiment with and apply the latest in machine learning technology to the tasks of catalog construction and analysis. SKICAT provides a unique environment for implementing these tools for any number of future scientific purposes. Initial scientific verification and performance tests have been made using galaxy counts and measurements of galaxy clustering from small subsets of the survey data, and a search for very high redshift quasars. All of the tests were successful, and produced new and interesting scientific results. Attachments to this report give detailed accounts of the technical aspects for multivariate statistical analysis of small and moderate-size data sets, called STATPROG. The package was tested extensively on a number of real scientific applications, and has produced real, published results.

Djorgovski, S. George↗

Multivariate Statistical Analysis Software Technologies for Astrophysical Research Involving Large Data Bases

We developed a package to process and analyze the data from the digital version of the Second Palomar Sky Survey. This system, called SKICAT, incorporates the latest in machine learning and expert systems software technology, in order to classify the detected objects objectively and uniformly, and facilitate handling of the enormous data sets from digital sky surveys and other sources. The system provides a powerful, integrated environment for the manipulation and scientific investigation of catalogs from virtually any source. It serves three principal functions: image catalog construction, catalog management, and catalog analysis. Through use of the GID3* Decision Tree artificial induction software, SKICAT automates the process of classifying objects within CCD and digitized plate images. To exploit these catalogs, the system also provides tools to merge them into a large, complex database which may be easily queried and modified when new data or better methods of calibrating or classifying become available. The most innovative feature of SKICAT is the facility it provides to experiment with and apply the latest in machine learning technology to the tasks of catalog construction and analysis. SKICAT provides a unique environment for implementing these tools for any number of future scientific purposes. Initial scientific verification and performance tests have been made using galaxy counts and measurements of galaxy clustering from small subsets of the survey data, and a search for very high redshift quasars. All of the tests were successful and produced new and interesting scientific results. Attachments to this report give detailed accounts of the technical aspects of the SKICAT system, and of some of the scientific results achieved to date. We also developed a user-friendly package for multivariate statistical analysis of small and moderate-size data sets, called STATPROG. The package was tested extensively on a number of real scientific applications and has produced real, published results.

Djorgovski, S. G.↗

Mercurian crater-filling classes constrain the emplacement process of the intercrater plains material

The multivariate method for statistical analysis of crater-filling classes, as presently applied to craters of more than 10-km diameter on Mercury, is noted to be superior to current techniques due to the obviation of any amalgamation of data from regions with different histories in order to arrive at the pertinent counting statistics. In this application to the Mercury intercrater plains, the process responsible for emplacement is constrained to be one that leaves the crater class distribution unaltered from that observed on the densely cratered terrain. The analysis explicitly eliminates the viability of the pulse emplacement of volcanics or basin ejecta.

Woronow, Alex↗

Regularized Differentiation for Bioburden Density Estimation in Planetary Protection

In this paper, we propose and investigate the performance of two novel shrinkage estimators for bioburden density estimation in planetary protection. The estimators are based on the regularized differentiation of a cumulative count of colony forming units collected throughout the data collecting session or the life cycle of the entire mission. The regularized differentiation recasts the problem of bioburden density estimation as a linear least squares problem. The least squares problem is then solved through regularization techniques, such as truncated singular value decomposition and penalized least squares. The regularization is necessary to avoid noise amplification during the differentiation of noisy data. The two regularization estimators are compared with four other commonly used estimators to simultaneously evaluate the means of multivariable independent Poisson distributions: the maximum likelihood, noninformative Bayes estimator with Jeffreys prior, Empirical Bayes using conjugate gamma-Poisson model with gamma parameters selected by method of moments, and the Clevenson-Zidek estimator. It is shown through computer-simulated data that the regularized differentiation based on ridge regression has the smallest mean-squared error among all estimators. The analysis of shrinkage mechanism implemented by regularized differentiation is performed, and it is shown that the regularized differentiation amounts to performing a weighted averaging of all the samples. The weights are determined by the regularization parameter automatically selected by the L-curve technique. Since the method of least squares makes no distributional assumptions about the data, it presents an attractive technique for bioburden density estimation when there are concerns about the misspecification of the distributional model. The paper concludes with the analysis of the bioburden data collected during InSight mission and directions for future work.

97 - MATHEMATICS AND COMPUTING↗