Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical graphics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

GENESPACE R Package (GENESPACE) v1.0

In short, the GENESPACE pipeline conducts analysis of orthology networks, constrained within syntenic regions. Since analyses are limited to local tests conducted within syntenic blocks, GENESPACE is agnostic to ploidy, duplicated regions, inversions or other whole-genome chromosomal complexities that are common across many evolutionary lineages. This advantage allows for evolutionary tests in polyploids (e.g. switchgrass, manuscript in review), species with ancient, but retained whole-genome duplications (e.g. pecan, manuscript in prep), high levels of tandem array proliferation (e.g. eukalypts, manuscript in review) and many other factors that can confound comparative genomic analyses. The major advances of GENESPACE are three-fold: First, this is the first R package to integrate visualization and analysis of large-scale comparative genomics. R, which offers a high-level environment for graphical and statistical exploration of data, is often speed- and memory-limited and not used for computationally intensive tasks such as comparative genomics. The highly efficient C++ scripts used in GENESPACE (via data.table) permit a much faster and computationally lightweight implementation of comparative genomics than is currently available. Second, the pipeline itself is novel. To the best of our knowledge, no other program accomplishes synteny-constrained and ploidy-agnostic comparative genomics. Since nearly all plants and many animals have a history of whole-genome duplications, this is a major and necessary advance to the field. Third, GENESPACE offers high-level and intuitive multi-genome graphical outputs. The dotplots and 'riparian' plots produced herein, which are produced entirely through original R code, are publication-ready and easily customizable.

Schmutz, Jeremy↗

BrazilClim : The overcoming of limitations of pre‐existing bioclimate data

Abstract Species distribution modelling has become instrumental in assessing the influence of environmental conditions on the occurrence or abundance of taxa. The set of environmental layers used for this purpose is a crucial aspect, for which different climate‐based (bioclimatic) datasets have been recently developed. These bioclimatic variables result from combinations of precipitation and temperatures surfaces. Here, we explored both the performance and possibility of improving some of the currently available bioclimatic databases, through an evaluation of the precipitation and temperatures surfaces used to generate them. For this purpose, we used a combination of statistic and graphic approaches. We focused on Brazil, not only due to its natural megadiversity, but also due to its continental size and orographic heterogeneity: an excellent ground for refining methods replicable elsewhere. We found a better match between the climatic data measured on‐field and Tropical Rainfall Measuring Mission (TRMM 3B43 v7) in the case of precipitation, and the surfaces provided by the National Oceanic and Atmospheric Administration (NOAA) in the case of temperatures, sources uncommonly used for species niche modelling. We gauge‐calibrated the best performing surfaces using machine‐learning algorithms and generated corrected surfaces that allowed us to create BrazilClim: a database of bioclimatic variables, based on improved primary surfaces, which will result in more assertive predicted distributions and more actual pictures of the species' ecological requirements for megadiverse Brazil, an approach replicable elsewhere. All primary and bioclimatic surfaces generated for this study may be freely downloaded.

Ramoni‐Perazzi, Paolo↗

EcoPLOT: dynamic analysis of biogeochemical data

Motivation: We have created EcoPLOT (parameterized linkage of omics-driven technologies), a web-app for the dynamic, interactive analysis of biogeochemical datasets that combines state-of-the-art analysis tools to statistically and graphically explore environmental, geochemical and microbiome datasets. Using the iterative random forest, a machine learning algorithm, EcoPLOT allows for the de novo discovery of drivers which exhibit significant impact on plant, microbial or soil dynamics. Availability and implementation: EcoPLOT is built entirely within the R language. It can be accessed through any system where R is installed, including Windows, Mac and most Linux systems. EcoPLOT is free to use and can be accessed at https://github.com/cdsanchez18/EcoPLOT.

59 BASIC BIOLOGICAL SCIENCES↗

Assessing correlated truncation errors in modern nucleon-nucleon potentials

We test the BUQEYE model of correlated effective field theory (EFT) truncation errors on Reinert, Krebs, and Epelbaum's semilocal momentum-space implementation of the chiral EFT (𝜒⁢EFT ) expansion of the nucleon-nucleon (NN) potential. This Bayesian model hypothesizes that dimensionless coefficient functions extracted from the order-by-order corrections to NN observables can be treated as draws from a Gaussian process (GP). We combine a variety of graphical and statistical diagnostics to assess when predicted observables have a 𝜒⁢EFT convergence pattern consistent with the hypothesized GP statistical model. Our conclusions are that, first, the BUQEYE model is generally applicable to the potential investigated here, which enables statistically principled estimates of the impact of higher EFT orders on observables. Second, parameters defining the extracted coefficients such as the expansion parameter 𝑄 must be well chosen for the coefficients to exhibit a regular convergence pattern—a property we exploit to obtain posterior distributions for such quantities. Third, the assumption of GP stationarity across lab energy and scattering angle is not generally met; this necessitates adjustments in future work. We provide a workflow and interpretive guide for our analysis framework, and show what can be inferred about probability distributions for 𝑄, the EFT breakdown scale Λ 𝑏 , the scale associated with soft physics in the 𝜒⁢EFT potential 𝑚 eff , and the GP hyperparameters. All our results can be reproduced using a publicly available Jupyter notebook, which can be straightforwardly modified to analyze other 𝜒⁢EFT NN potentials.

Bayesian methods↗

Collaborative Exploration of Scientific Datasets Using Immersive and Statistical Visualization: Preprint

We discuss the value of collaborative, immersive visualization for the exploration of scientific datasets and review techniques and tools that have been developed and deployed at the National Renewable Energy Laboratory (NREL). We believe that collaborative visualizations linking statistical interfaces and graphics on laptops and high-performance computing (HPC) with 3D visualizations on immersive displays (head-mounted displays and large-scale immersive environments) enable scientific workflows that further rapid exploration of large, high-dimensional datasets by teams of analysts. We present a framework, PlottyVR, that blends statistical tools, general-purpose programming environments, and simulation with 3D visualizations. To contextualize this framework, we propose a categorization and loose taxonomy of collaborative visualization and analysis techniques. Finally, we describe how scientists and engineers have adopted this framework to investigate large, complex datasets.

collaborative visualization↗

Collaborative Exploration of Scientific Datasets Using Immersive and Statistical Visualization

We discuss the value of collaborative, immersive visualization for the exploration of scientific datasets and review techniques and tools that have been developed and deployed at the National Renewable Energy Laboratory (NREL). We believe that collaborative visualizations linking statistical interfaces and graphics on laptops and high-performance computing (HPC) with 3D visualizations on immersive displays (head-mounted displays and large-scale immersive environments) enable scientific workflows that further rapid exploration of large, high-dimensional datasets by teams of analysts. We present a framework, PlottyVR, that blends statistical tools, general-purpose programming environments, and simulation with 3D visualizations. To contextualize this framework, we propose a categorization and loose taxonomy of collaborative visualization and analysis techniques. Finally, we describe how scientists and engineers have adopted this framework to investigate large, complex datasets.

collaborative visualization↗

How to Obtain the Redshift Distribution from Probabilistic Redshift Estimates

Abstract A reliable estimate of the redshift distribution n ( z ) is crucial for using weak gravitational lensing and large-scale structures of galaxy catalogs to study cosmology. Spectroscopic redshifts for the dim and numerous galaxies of next-generation weak-lensing surveys are expected to be unavailable, making photometric redshift (photo- z ) probability density functions (PDFs) the next best alternative for comprehensively encapsulating the nontrivial systematics affecting photo- z point estimation. The established stacked estimator of n ( z ) avoids reducing photo- z PDFs to point estimates but yields a systematically biased estimate of n ( z ) that worsens with a decreasing signal-to-noise ratio, the very regime where photo- z PDFs are most necessary. We introduce Cosmological Hierarchical Inference with Probabilistic Photometric Redshifts ( CHIPPR ), a statistically rigorous probabilistic graphical model of redshift-dependent photometry that correctly propagates the redshift uncertainty information beyond the best-fit estimator of n ( z ) produced by traditional procedures and is provably the only self-consistent way to recover n ( z ) from photo- z PDFs. We present the chippr prototype code, noting that the mathematically justifiable approach incurs computational cost. The CHIPPR approach is applicable to any one-point statistic of any random variable, provided the prior probability density used to produce the posteriors is explicitly known; if the prior is implicit, as may be the case for popular photo- z techniques, then the resulting posterior PDFs cannot be used for scientific inference. We therefore recommend that the photo- z community focus on developing methodologies that enable the recovery of photo- z likelihoods with support over all redshifts, either directly or via a known prior probability density.

79 ASTRONOMY AND ASTROPHYSICS↗

The circular bioeconomy: a driver for system integration

Background: Human and earth system modeling, traditionally centered on the interplay between the energy system and the atmosphere, are facing a paradigm shift. The Intergovernmental Panel on Climate Change’s mandate for comprehensive, cross-sectoral climate action emphasizes avoiding the vulnerabilities of narrow sectoral approaches. Our study explores the circular bioeconomy, highlighting the intricate interconnections among agriculture, forestry, aquaculture, technological advancements, and ecological recycling. Collectively, these sectors play a pivotal role in supplying essential resources to meet the food, material, and energy needs of a growing global population. We pose the pertinent question of what it takes to integrate these multifaceted sectors into a new era of holistic systems thinking and planning. Results: The foundation for discussion is provided by a novel graphical representation encompassing statistical data on food, materials, energy flows, and circularity. This representation aids in constructing an inventory of technological advancements and climate actions that have the potential to significantly reshape the structure and scale of the economic metabolism in the coming decades. In this context, the three dominant mega-trends—population dynamics, economic developments, and the climate crisis—compel us to address the potential consequences of the identified actions, all of which fall under the four categories of substitution, efficiency, sufficiency, and reliability measures. Substitution and efficiency measures currently dominate systems modeling. Including novel bio-based processes and circularity aspects might require only expanded system boundaries. Conversely, paradigm shifts in systems engineering are expected to center on sufficiency and reliability actions. Effectively assessing the impact of sufficiency measures will necessitate substantial progress in inter- and transdisciplinary collaboration, primarily due to their non-technological nature. In addition, placing emphasis on modeling the reliability and resilience of transformation pathways represents a distinct and emerging frontier that highlights the significance of an integrated network of networks. Conclusions: Existing and emerging circular bioeconomy practices can serve as prime examples of system integration. These practices facilitate the interconnection of complex biomass supply chain networks with other networks encompassing feedstock-independent renewable power, hydrogen, CO 2 , water, and other biotic, abiotic, and intangible resources. Elevating the prominence of these connectors will empower policymakers to steer the amplification of synergies and mitigation of tradeoffs among systems, sectors, and goals.

09 BIOMASS FUELS↗

Model Validation Database for Fires Involving Fuels at Liquefied Natural Gas Facilities

This document provides a description of the model evaluation protocol (MEP) database for fires involving liquefied natural gas (LNG) and processing fuels at LNG facilities. The purpose of the MEP is to provide procedures regarding the assessment of a model's suitability to predict thermal exclusion zones resulting from a fire. The database includes measurements from pool fire, jet fire, and fireball experiments which are provided in a spreadsheet. Users are to enter model results into the spreadsheet which automatically generates statistical performance measures and graphical comparisons with the experimental data. The intent of this document is to provide a description of the experiments and of the procedure required to carry out the validation portion of the MEP. In addition, the statistical performance measures, measurements for comparisons, and parameter variation are provided.

03 NATURAL GAS↗

Model Validation Database for Fires Involving Fuels at Liquefied Natural Gas Facilities (Version 2)

This document provides a description of the model evaluation protocol (MEP) database for fires involving liquefied natural gas (LNG) and processing fuels at LNG facilities. The purpose of the MEP is to provide procedures regarding the assessment of a model’s suitability to predict thermal exclusion zones resulting from a fire. The database includes measurements from pool fire, jet fire, and fireball experiments which are provided in a spreadsheet. Users are to enter model results into the spreadsheet which automatically generates statistical performance measures and graphical comparisons with the experimental data. The intent of this document is to provide a description of the experiments and of the procedure required to carry out the validation portion of the MEP. In addition, the statistical performance measures, measurements for comparisons, and parameter variation are provided.

03 NATURAL GAS↗

Kronecker-structured covariance models for multiway data

Many applications produce multiway data of exceedingly high dimension. Modeling such multi-way data is important in multichannel signal and video processing where sensors produce multi-indexed data, e.g. over spatial, frequency, and temporal dimensions. We will address the challenges of covariance representation of multiway data and review some of the progress in statistical modeling of multiway covariance over the past two decades, focusing on tensor-valued covariance models and their inference. We will illustrate through a space weather application: predicting the evolution of solar active regions over time.

97 MATHEMATICS AND COMPUTING↗

A Review of Bayesian Networks for Spatial Data

We report Bayesian networks are a popular class of multivariate probabilistic models as they allow for the translation of prior beliefs about conditional dependencies between variables to be easily encoded into their model structure. Due to their widespread usage, they are often applied to spatial data for inferring properties of the systems under study and also generating predictions for how these systems may behave in the future. We review published research on methodologies for representing spatial data with Bayesian networks and also summarize the application areas for which Bayesian networks are employed in the modeling of spatial data. We find that a wide variety of perspectives are taken, including a GIS-centric focus on efficiently generating geospatial predictions, a statistical focus on rigorously constructing graphical models controlling for spatial correlation, as well as a range of problem-specific heuristics for mitigating the effects of spatial correlation and dependency arising in spatial data analysis. Special attention is also paid to potential future directions for integration of Bayesian networks with spatial processes.

97 MATHEMATICS AND COMPUTING↗

Quantitative insights into the dislocation source behavior of twin boundaries suggest a new dislocation source mechanism

Pop-in statistics from nanoindentation with spherical indenters are used to determine the stress required to activate dislocation sources in twin boundaries (TBs) in copper and its alloys. The TB source activation stress is smaller than that needed for bulk single crystals, irrespective of the indenter size, dislocation density and stacking fault energy. Because an array of pre-existing Frank partial dislocations is present at a TB, we propose that dislocation emission from the TB occurs by the Frank partials splitting into Shockley partials moving along the TB plane and perfect lattice dislocations, both of which are mobile. The proposed mechanism is supported by recent high resolution transmission electron microscopy images in deformed nanotwinned (NT) metals and may help to explain some of the superior properties of nanotwinned metals (e.g. high strength and good ductility), as well as the process of detwinning by the collective formation and motion of Shockley partial dislocations along TBs.

36 MATERIALS SCIENCE↗

A geospatial risk analysis graphical user interface for identifying hazardous chemical emission sources

Background: Performing back trajectory and forward trajectory using the Hybrid Single-Particle Lagrangian Integrated Trajectory Model (HYSPLIT) is a reliable approach for assessing particle transport after release among mid-field atmospheric models. HYSPLIT has an externally facing online interface that allows non-expert users to run the model trajectories without requiring extensive training or programming. However, the existing HYSPLIT interface is limited if simulations have a large amount of meteorological data and timesteps that are not coincident. The objective of this study is to design and develop a more robust tool to rapidly evaluate hazard transport conditions and to perform risk analysis, while still maintaining an intuitive and user-friendly interface. Methods: HYSPLIT calculates forward and backward trajectories of particles based on wind speed, wind direction, and the corresponding location, timestamp, and Pasquill stability classes of the regions of the atmosphere in terms of the wind speed, the amount of solar radiation, and the fractional cloud cover. The computed particle transport trajectories, combined with the online Proton Transfer Reaction-Mass Spectrometry (PTR-MS) data (https://figshare.com/articles/dataset/ARL_Data_from_PROS_station_at_Hanford_site/19993964), can be used to identify and quantify the sources and affected area of the hazardous chemicals’ emission using the potential source distribution function (PSDF). PSDF is an improved statistical function based on the well-known potential source contribution function (PSCF) in establishing the air pollutant source and receptor relationship. Performing this analysis requires a range of meteorological and pollutant concentration measurements to be statistically meaningful. The existing HYSPLIT graphical user interface (GUI) does not easily permit computations of trajectories of a dataset of meteorological data in high temporal frequency. To improve the performance of HYSPLIT computations from a large dataset and enhance risk analysis of the accidental release of material at risk, a geospatial risk analysis tool (GRAT-GUI) is created to allow large data sets to be processed instantaneously and to provide ease of visualization. Results: The GRAT-GUI is a native desktop-based application and can be run in any Windows 10 system without any internet access requirements, thus providing a secure way to process large meteorological datasets even on a standalone computer. GRAT-GUI has features to import, integrate, and convert meteorological data with various formats for hazardous chemical emission source identification and risk analysis as a self-explanatory user interface. The tool is available at https://figshare.com/articles/software/GRAT/19426742.

97 MATHEMATICS AND COMPUTING↗

SF_Microbe_Methane (SFMM) v1

Scripts for analyzing microbial community taxonomy and function, running statistical tests, and making complex graphics.

de Mesquita, CliftonB↗

On reading Youden: Learning about the practice of statistics and applied statistical research from a master applied statistician

From reading William John “Jack” Youden’s books and articles, Youden (1900-1971), an analytical chemist, becomes an applied statistician by the time he joins the National Bureau of Standards (NBS) in 1948. Here, this article traces his transition from chemist to applied statistician and what his body of work mostly at NBS (1948-1965) demonstrates about his practice of statistics and the role that applied statistical research plays in it. There is much we can learn from a master applied statistician.

42 ENGINEERING↗

A quantitative comparison of the fingerprint of twinned microstructures through surface and three-dimensional techniques

Assessing the fingerprint of a material’s microstructure is key for supporting materials design. With the emergence of a wide range of 3D characterization techniques, it is critical to understand the main differences in fingerprints reconstructed from 2D and 3D datasets. To this end, we introduce a graph-based microstructure reconstruction framework that enables structural comparisons of twin domain networks in high purity Ti using 3D and 2D electron backscatter diffraction. Insights into the structure of the twin networks are facilitated by combining statistical analysis of twin crystallography with visual and graphical analysis of the novel graph abstractions of the twins. We demonstrate that compared to 3D reconstructions, conventional 2D views of twinning miss key aspects of the microstructure including the high interconnectivity of domains into networks that span the full reconstruction volume. The reduced cross-grain and in-grain twin connectivity typically observed in 2D has notable implications on our understanding of how twinning mediates the plastic response of microstructures and how twin networks evolve. It is thus clear that 3D characterization is critical for accurately inferring both twin network morphologies as well as the key unit processes facilitating network formation.

36 MATERIALS SCIENCE↗

Tractable minor-free generalization of planar zero-field Ising models

In this work, we present a new family of zero-field Ising models over N binary variables/spins obtained by consecutive 'gluing' of planar and O(1)-sized components and subsets of at most three vertices into a tree. The polynomial time algorithm of the dynamic programming type for solving exact inference (computing partition function) and exact sampling (generating i.i.d. samples) consists of sequential application of an efficient (for planar) or brute-force (for O(1)-sized) inference and sampling to the components as a black box. To illustrate the utility of the new family of tractable graphical models, we first build a polynomial algorithm for inference and sampling of zero-field Ising models over K 33 -minor-free topologies and over K 5 -minor-free topologies—both of which are extensions of the planar zero-field Ising models—which are neither genus- nor treewidth-bounded. Second, we empirically demonstrate an improvement in the approximation quality of the NP-hard problem of inference over the square-grid Ising model in a node-dependent nonzero 'magnetic' field.

97 MATHEMATICS AND COMPUTING↗