Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Kinetic, metabolic, and statistical analytics: addressing metabolic transport limitations among organelles and microbial communities

Microbial organisms engage in a variety of metabolic interactions. A crucial part of these interactions is the exchange of molecules between different organelles, cells, and the environment. The main forces mediating this metabolic exchange are transporters. This transport can be difficult to measure experimentally because several transport mechanisms remain opaque. However, theoretical calculations about the inputs and outputs of cells via metabolic exchanges have enabled the successful inference of the workings of intra-organismal and inter-organismal systems. Kinetic, metabolic, and statistical modeling approaches in combination with omics data are enhancing our knowledge and understanding about metabolic exchange and mass resource allocation. Furthermore, this model-driven analytics approach can guide effective experimental design and yield new insights into biological function and control.

59 BASIC BIOLOGICAL SCIENCES↗

Characterizing the Uncertainty of Measurement of Traceable Isotope Ratios with Bayesian Statistical Techniques

Analytical techniques such as multicollector—inductively coupled plasma—mass spectrometry (MC-ICP-MS) are routinely employed at SRNL, other National Laboratories, and in academia to determine the precise isotopic composition of diverse natural and anthropogenic samples (e.g., rocks and nuclear materials). Quantifying and reporting uncertainty in such analyses, while regularly performed, have a rigorous statistical foundation. The Guide to the Expression of Uncertainty in Measurement 4 (GUM) outlines conventional techniques used to assess such uncertainty. As the accessibility and speed of statistical computing increase, there is a need to modernize conventional techniques. For example, Supplement 1 to the 3rd to the GUM suggests the use of approximation methods as an updated approach to the GUM.

McLarty, Ellis C.↗

In situ feature analysis for large-scale multiphase flow simulations

The study of multiphase flow is essential for designing chemical reactors such as fluidized bed reactors (FBR), as a detailed understanding of hydrodynamics is critical for optimizing reactor performance and stability. An FBR allows scientists to conduct different types of chemical reactions involving multiphase materials, especially interaction between gas and solids. During such complex chemical processes, the formation of void regions in the reactor, generally termed as bubbles, is an important phenomenon. The study of these bubbles has a deep implication in predicting the reactor’s overall efficiency. But physical experiments needed to understand bubble dynamics are costly and non-trivial due to the technical difficulties involved and harsh working conditions of the reactors. Therefore, to study such chemical processes and bubble dynamics, a state-of-the-art computational simulation MFIX-Exa is being developed. Despite the proven accuracy of MFIX-Exa in modeling bubbling phenomena, the large-scale output data prohibits the use of traditional post hoc analysis capabilities in both storage and I/O time. Herein, to address these issues and allow the application scientists to explore the bubble dynamics in an efficient and timely manner, we have developed an end-to-end analytics pipeline that enables in situ detection of bubbles, followed by a flexible post hoc visual exploration methodology of bubble dynamics. The proposed method enables interactive analysis of bubbles, along with quantification of several bubble characteristics, enabling experts to understand the bubble interactions in detail. Positive feedback from the experts has indicated the efficacy of the proposed approach for exploring bubble dynamics in very-large-scale multiphase flow simulations.

97 MATHEMATICS AND COMPUTING↗

A Probabilistic Model-Based Diagnostic Framework for Nuclear Engineering Systems

A fault diagnostic framework was investigated in this study for applications in thermal–hydraulic systems of nuclear power plants. The proposed framework consists of quantitative model-based diagnosis, statistical change detection and probabilistic reasoning. The use of physics-based diagnostic models provides high detection sensitivity and allows noise and measurement uncertainty to be incorporated robustly. Performance-related parametric models for each component are constructed based on first principles. Numerical model residuals are generated using the concept of analytical redundancy. Statistical change detection methods are employed to detect non-zero residuals in the presence of uncertainty. The diagnosis task is performed using Bayesian inference to detect and localize possible faults. Application to a single-phase heat exchanger for demonstration showed that the proposed probabilistic framework can provide improved results in comparison with traditional approaches while remaining less sensitive to false alarms in the presence of measurement and modeling uncertainty.

Bayesian network↗

Distribution Development for Residual Inventory at the New York West Valley Site - 20400

The New York State Energy Research and Development Authority (NYSERDA) is the owner of the Western New York Nuclear Service Center (WNYNSC), a 1,351 ha site located approximately 48 km south of Buffalo, New York. In 1962, Nuclear Fuel Services, Inc. (NFS) entered into Agreements with the Atomic Energy Commission and New York State to construct the first commercial reprocessing plant of nuclear fuel in the United States. NFS, a private company, built and operated the spent fuel reprocessing plant and waste disposal facilities, processing 640 Mg of spent nuclear fuel from 1966 to 1972 under an Atomic Energy Commission license. Nuclear fuel reprocessing operations ended in 1972 and never reopened, leaving behind radioactive and chemical wastes. Operations led to contamination in a number of facilities and locations. Some of that contamination has migrated from waste disposal zones to other layers, formations, and features on and off the WNYNSC. Phase I decommissioning activities are ongoing and involve the removal of a number of areas and structures that have been associated with contamination. The purpose of this work is to outline the approach for characterizing contamination not associated with disposed wastes, contaminated structures, or specific releases. In this work, the term, residual radiological activity, is used to describe environmental contamination that exists subsequent to the completion of Phase I decommissioning activities, that is not associated with disposed wastes, contaminated structures, or specific releases. Contamination from the Site was quantified relative to data that characterize the concentrations of radionuclides that exist in background. Background concentrations are those present in the area but having no influence from Site related activities. The existence of residual radiological activity that is elevated relative to background has the potential to contribute to future risks to human health and the environment. As a consequence, the residual inventory information is used to inform the West Valley Probabilistic Performance Assessment (PPA) model to characterize potential future risks to human health and the environment. The centralized West Valley Data Management System (DMS) was the source of information for the data assembled in this analysis. The DMS is a fairly large compilation consisting of thousands of records from investigation studies, with sample dates ranging from 1990 to present. Samples from monitoring wells, boreholes, geoprobe studies, surface water, surface soils, storm water outfalls, ventilation stack filters, plant and animal tissues, and more are included in the DMS. Results are typically reported in units of activity per unit volume. For the purpose of the analyses presented here, all results were converted into consistent units of pCi per unit volume. Since 1990, data have been collected from various locations across the WNYNSC at different times with varying frequency over the course of several decades. As a consequence, a number of potential issues can arise with respect to the assembly of a dataset that is deemed adequate for the characterization of residual radiological activity. These issues were assessed and resolved to the extent possible through careful consideration of the properties of the distributions. The intent was to use data which characterize the current state of the Site. Radionuclides can be designated to one of several groups depending on their origin. In this work the groups considered were 1) Naturally Occurring Radioactive Material (NORM), 2) fallout, and 3) Other (including power plant, medical research, etc). This grouping is a useful construct with respect to the interpretation of fixed laboratory results. For example, NORM radionuclides that exist within a decay chain should have approximately equivalent distributions of concentrations if they are representative of background conditions. Insights such as these can be used as a check to identify sample results that need to be further investigated or omitted due to issues associated with reported results from fixed laboratory analyses. This type of analysis provided a foundation for the assessment of the adequacy of sample results for use in subsequent components of an assessment. The general process for the assessment of residual radiological contamination at the Site consists of a sequence of several steps. First, for each analyte, several statistical tests were performed to assess the weight of evidence against the null hypothesis that the mean of the distribution of concentrations was equal to zero. If the mean of the distribution of concentrations for a given radionuclide was not found to be greater than zero, then it was removed from consideration as a component of the residual radiological contamination. If there was significant evidence to reject the hypothesis of the mean being equal to zero, the second step was to compare the distribution of the data from the Site to that of the corresponding background. A suite of tests was used to compare the distributions of the site and background data. The results of these tests were collectively used to determine if site data are elevated relative to background. The third step was to develop distributions using a Bayesian framework to characterize the distribution of mean of the increment present above background for each of the radionuclides. The Bayesian model implemented allowed for the comparison of site-specific records to background concentrations to better approximate contamination attributed to the Site. A final screening step was employed for radionuclides that exceed background. This screening step compared 95% upper confidence limits (UCLs) from the increment distribution developed in the previous step to the risk screening levels. This approach yields a list of analytes that were determined to be elevated relative to background.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Multi-omics Characterization of the Host Response to COVID-19

This project is a multi-disciplinary collaboration between investigators at PNNL with expertise in mass spectrometry (MS)-based omics technology development, omics measurement methods development and application, statistics, machine learning and integration of disparate datasets for a systems-level understanding, and expertise in pathogenic coronaviruses, and investigators at the University of Wisconsin-Madison (UW-Madison) with expertise in pathogenic respiratory viruses (e.g. influenza). The goal of this project is to obtain a comprehensive picture of the human host factors critical for the outcome of SARS-CoV-2 infection. We will generate broad untargeted multi-omics profiles using both state-of-the-art and novel instrumentation and approaches to enable the identification of the molecular mechanisms and host response pathways that impact human COVID-19 outcomes. We anticipate these results will lead to the generation of biomarker panels that are predictive of disease outcomes and mechanistic hypotheses that can be further interrogated in future studies and will provide the basis for vaccine or therapeutic development. To do so, we are obtaining and analyzing blood samples from COVID-19 patients with a range of disease outcomes that were treated at the Center Hospital of the National Center for Global Health and Medicine in Tokyo, Japan and other collaborating hospitals in our network. Specifically, this project will fund proteomics and metabolomics analyses of clinical COVID samples, machine learning-based integration of the data, and pathway-based interpretation of the data. This project was funded in June 2020. In the time span of June to September 2020, the project team developed an analytically and statistically robust analysis plan and made various preparations to facilitate sample receipt from our UW-Madison collaborators. This included blocking and randomization of sample prep orders, ordering of reagents and reference materials, and shipping of materials needed for preparation of the samples under BSL3 conditions to our collaborators at UW-Madison. As of FY21, this project has been picked up via a sponsor, the Naval Medical Research Center, which will cover the remainder of the proposed scope of work.

60 APPLIED LIFE SCIENCES↗

Machine learning to discover mineral trapping signatures due to CO 2 injection

Mineral trapping is pursued as a geological CO 2 sequestration (GCS) mechanism because it permanently stores CO 2 in solid phases or minerals. However, CO 2 mineral-trapping mechanisms are poorly understood due to (1) lack of sufficient field and laboratory data characterizing these complex processes, and (2) challenges to develop site-specific reactive-transport models coupling fluid flow and geochemical reactions occurring at various temporal (from milliseconds to years) and spatial (from pore (millimeters) to field (kilometers)) scales. Reactive transport with additional complexities such as heterogeneity can make the simulation outputs even more difficult to interpret because of complex nonlinearity and multi-scale interdependencies. Furthermore, the values of model outputs such as concentrations can vary by several orders of magnitude, making it harder to correlate and characterize the impact of the variables via traditional data interpretation techniques such as exploratory data analyses. Recently, machine learning (ML) has shown promise in feature discovery and in highlighting hidden mechanisms that cannot be obtained by existing data-analytics and statistical methods. In this study, we applied an unsupervised ML approach, non-negative matrix factorization with custom -means clustering (NMF) to the data generated by reactive-transport simulations of GCS. The reactive-transport data consisted of 19 attributes, including four physio-chemical variables (pH, porosity, aqueous CO 2 , and sequestered CO 2 ), six chemical species (K + , Na + , HCO, Ca 2+ , Mg 2+ , Fe 2+ ), and four carbonate minerals (calcite, dolomite, siderite, and ankerite), a feldspar mineral (albite), and four clay minerals (illite, clinochlore, kaolinite, and smectite) over a period of 200 years of simulation time. Furthermore, the simulation data used was for Morrow B sandstone at the Farnsworth hydrocarbon unit in Texas. Data are sampled at two locations within the model domain: (1) at the injection well and (2) 200 m west of the injection well. The injection was performed for a period of 10 years. Using NMF, we estimated the temporal interdependencies among the 19 attributes over a span of 200 years. We found that NMF was able to identify four reaction stages and their dominant attributes; these cannot be directly discerned through traditional visualization (e.g., line plots, Pareto analysis, Glyph-based visualization methods) or exploratory data analysis tools of the simulation data. The four stages were: reactions in the injection phase followed by short-, mid-, and long-term reactions. The NMF analysis also revealed that 10 among the 19 attributes are dominant. These dominant attributes for mineral trapping include calcite, dolomite at injection well, siderite at 200 m away from the injection well, clinochlore, kaolinite, Na + , K + , Ca 2+ , Mg 2+ , pH, and aqeuous CO 2 . Finally, at late times (65–200 years), our results showed that calcite plays a major role in mineral trapping with insignificant contribution from siderite, ankerite, and clay minerals. These findings make the proposed unsupervised ML-model attractive for reactive-transport sensing towards real-time GCS monitoring.

54 ENVIRONMENTAL SCIENCES↗

The Completed SDSS-IV Extended Baryon Oscillation Spectroscopic Survey: N-body Mock Challenge for Galaxy Clustering Measurements

We develop a series of N-body data challenges, functional to the final analysis of the extended Baryon Oscillation Spectroscopic Survey (eBOSS) Data Release 16 (DR16) galaxy sample. The challenges are primarily based on high-fidelity catalogues constructed from the Outer Rim simulation - a large box size realization (3h(-1) Gpc) characterized by an unprecedented combination of volume and mass resolution, down to 1.85 x 10(9) h(-1)M(circle dot). We generate synthetic galaxy mocks by populating Outer Rim haloes with a variety of halo occupation distribution (HOD) schemes of increasing complexity, spanning different redshift intervals. We then assess the performance of three complementary redshift space distortion (RSD) models in configuration and Fourier space, adopted for the analysis of the complete DR16 eBOSS sample of Luminous Red Galaxies (LRG5). We find all the methods mutually consistent, with comparable systematic errors on the Alcock-Paczynski parameters and the growth of structure, and robust to different HOD prescriptions - thus validating the robustness of the models and the pipelines used for the baryon acoustic oscillation (BAO) and full shape clustering analysis. In particular, all the techniques are able to recover and alpha(11) to within 0.9 per cent, and f sigma(8) to within 1.5 per cent. As a by-product of our work, we are also able to gain interesting insights on the galaxy-halo connection. Our study is relevant for the final eBOSS DR16 'consensus cosmology', as the systematic error budget is informed by testing the results of analyses against these high-resolution mocks. In addition, it is also useful for future large-volume surveys, since similar mock-making techniques and systematic corrections can be readily extended to model for instance the Dark Energy Spectroscopic Instrument (DESI) galaxy sample.

cosmology: theory, large-scale structure of Univer↗

Advanced Health Information Technology Analytic Framework and Application to Hazard Detection

Health Information Technology (HIT) aims to improve healthcare outcomes by organizing and analyzing various health-related data. With data accumulating at a staggering rate, the importance of real-time analytics has been increasing dramatically, shifting the focus of informatics from batch processing to streaming analytics. HIT is also facing unprecedented challenges in adapting to this new requirement and leveraging advanced IT technologies. This paper introduces a HIT data and compute platform that supports multi-granularity real-time analytics from heterogeneous data sources. The paper first identifies functional requirements and proposes a framework that satisfies the requirements using state-of-the-art big data technologies including Apache Kafka, Spark Structured Streaming Engine, and Delta Lake. To demonstrate its capability to support data analytics in multiple time granularities analytics, a statistical process control-based hazard detection algorithm has been implemented on top of the framework to detect unexpected hazards from order cancellation data of the Department of US Veterans Affairs (VA) in near real-time.

Kumar, Mohit↗

An analytic approximation to the covariance between pre- and post-reconstruction galaxy two-point statistics

We present a simple analytic approximation for the covariance between pre-reconstruction galaxy power spectrum measurements and post-reconstruction two-point correlation functions. This cross-covariance is essential for joint analyses that combine full-shape clustering information with baryon acoustic oscillation (BAO) measurements, as commonly performed in modern spectroscopic surveys. Our model builds on the disconnected contribution to the covariance and accounts for the damping of correlations due to the BAO reconstruction process. We validate our analytic prescription against numerical simulations from the Dark Energy Spectroscopic Instrument (DESI), testing both idealized cubic geometries and realistic survey configurations including complex footprints and fiber assignment effects. Despite neglecting survey window functions in the analytic calculation, we find excellent agreement with simulation-based covariances and demonstrate that cosmological parameter constraints are virtually unchanged when using our approximation. Our results show that the pre-post cross-covariance is sufficiently small that even approximate treatments are adequate for cosmological inference, opening a pathway toward fully analytic covariance matrices for next-generation galaxy surveys.

baryon acoustic oscillations↗

Statistical uncertainties of the v n { 2 k } harmonics from Q cumulants

Analytic formulas to calculate statistical uncertainties of v n { 2 k } harmonics extracted from the Q cumulants are presented. The Q cumulants are multivariate polynomial functions of the weighted means of 2m-particle azimuthal correlations, $\langle\langle 2 m \rangle\rangle$. Variances and covariances of the $\langle\langle 2 m \rangle\rangle$ are included in the analytic formulas of the uncertainties that can be calculated simultaneously with the calculations of the v n { 2 k } harmonics. The calculations are performed using a simple toy model, which roughly simulates elliptic flow azimuthal anisotropy with magnitudes around 0.05. The results are compared with the results obtained by the many data subsets, and by the bootstrapping method. The first one is a common way of estimation of the statistical uncertainties of the v n { 2 k } harmonics in a real experiment. In order to increase precision in the measurement of the v n { 2 k } harmonics, a large number of 15 000 subsets and the same number of the resampling in the bootstrap method is used. Unlike the other ways of the analytic calculation of the v n { 2 k } statistical uncertainties, our proposal that includes the use of squared weights in the calculation of both the variances and covariances, gives the best agreement with the results obtained from the subsets and bootstrap method. Additionally, a recurrence equation between Q cumulants of any order is also presented.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Installation Assignment Durations and Patterns of Army Personnel

The purpose of this document is to report on the development of personnel assignment duration statistical metrics and statistical descriptions for groups of Army personnel aggregated by rank, career management field, and other descriptors for 32 Continental United States (CONUS) Army installations. This information can be used for generating evidence-based measures of personnel exposure times associated with environmental health considerations at the installation and career management field levels. Assignment duration statistical descriptions were developed to provide temporal metrics for groups of personnel at Army installations as foundational data to support a variety of assessments. Statistical results of personnel assignment durations at the installation level are provided in several formats. The 50th percentile assignment duration represents the median number of days personnel in groups aggregated by rank and career management field spend at the 32 CONUS Army installations in this study. Likewise, other duration percentiles (90th) and quartiles of assignment durations for groups aggregated by rank and career management field spend at each installation were developed. Assignment path analysis provide statistical descriptions of assignment durations at Army installations throughout increments of time (e.g., 5 years, 10 years). Key groupings of assignment paths were achieved by analyzing the duration of assignments at installations in chronological order from entry throughout specified increments of military service to generate the most common assignment paths across the 32 CONUS Army installations in this study. The developed assignment duration analytics and duration statistical descriptions provide foundational temporal metrics for groups of personnel at 32 CONUS Army installations to support a variety of assessments including those by human resource planners and environmental health specialists.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Application of Systems Engineering Principles and Techniques in Biological Big Data Analytics: A Review

In the past few decades, we have witnessed tremendous advancements in biology, life sciences and healthcare. These advancements are due in no small part to the big data made available by various high-throughput technologies, the ever-advancing computing power, and the algorithmic advancements in machine learning. Specifically, big data analytics such as statistical and machine learning has become an essential tool in these rapidly developing fields. As a result, the subject has drawn increased attention and many review papers have been published in just the past few years on the subject. Different from all existing reviews, this work focuses on the application of systems, engineering principles and techniques in addressing some of the common challenges in big data analytics for biological, biomedical and healthcare applications. Specifically, this review focuses on the following three key areas in biological big data analytics where systems engineering principles and techniques have been playing important roles: the principle of parsimony in addressing overfitting, the dynamic analysis of biological data, and the role of domain knowledge in biological data analytics.

dynamic analysis↗