Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hypothesis tests”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

VAINE: Visualization and AI for Natural Experiments

Natural experiments are observational studies where the assignment of treatment conditions to different populations occur by chance ``in the wild''. Researchers from fields such as economics, healthcare, and the social sciences leverage natural experiments to conduct hypothesis testing and causal effect estimation for treatment and outcome variables that would otherwise be costly, infeasible, or unethical. In this paper, we introduce VAINE (Visualization and AI for Natural Experiments), a visual analytics tool for identifying and understanding natural experiments from observational data. We then demonstrate how VAINE can be used to validate causal relationships, estimate average treatment effects, and identify statistical phenomena such as Simpson’s paradox through two use cases.

Guo, Grace↗

Tissue scale agent-based simulation of premalignant progressions in Barrett’s esophagus

Barrett’s esophagus (BE) is a benign condition of the distal esophagus that initiates a multistage pathway to esophageal adenocarcinoma (EAC). Short of frequent intrusive (and costly) surveillance, effective screening for neoplasia in BE populations is yet to be established since progressors are rare and virtually undetectable without routine biopsies, which often sample only a small portion of the BE tissue. As a result, reliable estimation of the true prevalence of dysplasia in a BE population and evidence-based optimization of screening for at-risk individuals is challenging. Data-driven microsimulations, i.e., model-generated instances of disease history in a predefined virtual population, have found utility in the EAC screening literature as low-overhead alternatives to real-world hypothesis testing of optimal interventions for dysplasia. Despite the successes, computational limitations, paucity of knowledge and data on Barrett’s dysplasia, and the complexities of disease progression as a multiscale multiphysics process have hindered the treatment of disease progression in BE as a spatial process. Agent-based modeling of nucleation and proliferation processes in dysplasia warrants exploration in this context as an approximation that operates at a trade-off between computational tractability and precise representation of the composition and physics of the substrate (tissue). In this study, we describe spatially resolved simulations of premalignant progression toward EAC in a coarse-grained model of Barrett’s tissue that resolves the metaplastic tissue at a length scale of 0.42 mm (~3300 crypts/mm 2 ). Finally, the model is calibrated to reproduce historical high-grade dysplasia prevalence when model-generated patients are screened using the Seattle protocol.

59 BASIC BIOLOGICAL SCIENCES↗

Covariate Dependent Sparse Functional Data Analysis

This study proposes a method to incorporate covariate information into sparse functional data analysis. The method aims at cases where each subject has a limited number of longitudinal measurements and is associated with static covariates. This research is motivated by several use cases in practice. One representative example is void swelling, a nuclear-specific material degradation mechanism. Void swelling is affected by many covariates, including alloy composition and irradiation type. How to accurately model the complicated joint effects of such covariates on the swelling process is the key to mitigating the effect of swelling and ensuring safe operation. Unlike most of the existing methods, the proposed method can handle high-dimensional covariates with the informative covariate identification procedure and sparse and irregularly spaced measurements, that is, does not require complete or dense observations. The main innovation of the proposed method is that we model the variation coming from covariates and the variation left conditioned on covariates, such that the functional principal component analysis and Gaussian process can be conducted in a unified manner. Further, we also propose a systematic approach to identify important covariates in the hypothesis testing context. The methodology is demonstrated on applications in nuclear engineering and healthcare and simulation studies.

42 ENGINEERING↗

Model Inputs, Outputs, and Scripts associated with: “Spatial microbial respiration variations in the hyporheic zones within the Columbia River Basin”

This data package is associated with the publication “Spatial microbial respiration variations in the hyporheic zones within the Columbia River Basin” published in the Journal of Geophysical Research: Biogeosciences (Son et al. 2022) available at doi: 10.1029/2021JG006654. This data package includes the key model inputs/outputs of the river corridor model for the Columbia River Basin (CRB) and the model source codes, which were used in the manuscript. The model is a carbon-nitrogen-coupled river corridor model (RCM), and the model is used to quantify hyporheic zone (HZ) aerobic and anaerobic respiration at the NHDPLUS stream reach scales. The RCM used in this study combines empirical substrate models derived from observations and three microbially driven reactions to compute respiration of the HZ for each National Hydrography Dataset (NHD) reach within the CRB. The reactions in HZs of each NHD reach include anaerobic respiration and two-step anaerobic respiration via denitrification. Our HZ respiration estimates are limited to the lotic (or flowing) stream/river systems, and do not account for the respiration process in water column. Note that the RCM only simulates the HZ’s contribution to the dissolved carbon dioxide (CO2) concentrations in the streams, and the CO2 emissions to the atmosphere are not modelled. The model computes at hourly timesteps because of the fast reaction rates. The key input data of the model are exchange flux, residence time, and stream solute (dissolved organic carbon (DOC), dissolved oxygen (DO), and nitrate concentrations). These inputs are constant over time and represent long-term averaged values.This modeling framework successfully quantified HZ respiration components over multiple scales. It revealed key mechanisms driving the spatial variation of HZ aerobic and anaerobic respiration in reaches with varying hydrologic and substrate conditions. Thus, this modeling study offers a testing hypothesis in different river system (e.g., climate and biomes) for the HZ respiration processes, and can be used as a sampling design tool for large-scale HZ experimental studies.This dataset contains five folders: (1) model_inputs, (2) model_outputs, (3) Rscripts, (4) figures, and (5) model_codes. It also contains a readme, file level metadata (FLMD), and data dictionary (dd). Please see the FLMD for a list of all the files contained in this data package and descriptions for each. The model_inputs folder contains the model inputs used to drive the model simulations. The model_outputs folder contains key model output files from the river corridor model. The Rscripts folder contains the Rscripts for pre- and post- processing model results. The figures folder contains the raw figures associated with the manuscript. The model_codes folder includes key model source codes/input files. All files are .jpg, .jpeg, .out, .e, .od, .dat, .sub, .F90, .0, .R, .sbx, .cpg, .sbn, .shx, .shp, .dbf, .prj, .tfw, .tif, .xml, .pdf, or .csv.

54 ENVIRONMENTAL SCIENCES↗

Forecasting generative amplification

Generative networks are perfect tools to enhance the speed and precision of LHC simulations. Especially when generating events beyond the size of the training dataset, it is important to understand their statistical precision. We present two complementary methods to estimate the amplification factor without large holdout datasets. Averaging amplification uses Bayesian networks or ensembling to estimate amplification from the precision of integrals over given phase-space volumes. Differential amplification uses hypothesis testing to quantify amplification without any resolution loss. Applied to state-of-the-art event generators, both methods indicate that amplification is already possible in specific regions of phase space.

Bahl, Henning [Heidelberg Univ. (Germany)] (ORCID:↗

Systems Analysis of the Physiological and Molecular Mechanisms of Sorghum Nitrogen Use Efficiency, Water Use Efficiency and Interactions with the Soil Microbiome (Final Report for DE-SC0014395)

The specific project objectives were to: 1) Conduct deep census surveys of root microbiomes concurrent with phenotypic characterizations of a diverse panel of sorghum genotypes across multiple years to define the microbes associated with the most productive lines under drought and low nitrogen conditions. 2) Associate systems-level genotypic, microbial, and environmental factors with improved sorghum performance using robust statistical approaches. 3) Develop culture collections of sorghum root/leaf associated microbes that recapitulate root-enriched sequences defined in the census. 4) Perform controlled environment experiments for in-depth characterization and hypothesis testing of G sorghum x G microbe x E interactions . Validate physiological mechanisms, map genetic loci for stress tolerance, and determine the persistence of optimal microbial strains under greenhouse and field conditions.

59 BASIC BIOLOGICAL SCIENCES↗

3P Program: Phenotyping X Prediction = Productivity (Final Scientific/Technical Report)

The goal of the 3P Program was to establish integrated, real-time phenotyping and to analyze above- and below-ground plant architecture and total carbon partitioning and allocation to predict heterosis and develop superior crop hybrids by fully leveraging the Sorghum gene pool. There were two overarching themes: 1) the development of a new crop improvement approach utilizing advances in high-throughput phenotyping (HTP), computing, and genomics for public dissemination and 2) leveraging this platform for sorghum crop improvement and commercialization. The Clemson team worked on creating genomic resources and using both statistical learning and high-throughput phenotyping in genomics-assisted breeding. Research was broadly interested in the genetics of carbon partitioning, with the aim of improving crop performance and achieving sustainability. The technology and resources created can be readily found in the public domain and serve to advance scientific understanding of crop genomics and breeding. Genomic prediction was able to identify top crosses to be made, and a hybrid prediction pipeline is in place to drive year-over-year genetic gain. Roots have long been ignored by plant breeders and agronomists, not because they are unimportant but because they are hard to measure. This is an untapped white space of potential insight and innovation. To address this, Hi Fidelity Genetics developed the RootTracker to measure roots in the field on a continuous basis. A database system called RootTracker Tracker was developed to handle data coming from the RootTrackers. In using this device, valuable data was observed for plant breeding, hydrochemical development, and other agricultural biology applications. Carnegie Mellon’s goal was developing new techniques to generate high-resolution 3D models of plants from data collected in the field. The idea was that more useful and more informative phenotypes could be extracted by resolving small features, such as seeds and flowers, and that by modeling in 3D, the spatial structure of plants could be examined. To achieve this, multiple images collected by a new small format structured light stereo imager were fused together. A sorghum panicle modeling pipeline was developed to allow the collection and processing of data. Carolina Seed Systems is an agricultural technology company focused on decarbonizing the agricultural system. Their technology pipeline serves to drive fundamental progress towards creation and distribution of carbon negative crops. The genomic and the engineering technology developed through the 3P Program was leveraged to deliver both value and sustainability from the grower to the consumer. Promising sorghum hybrids were scaled up and commercialized. The overall goal of our research was to integrate, create, and deploy genetic and engineering concepts and technologies to enhance crop productivity in a sustainable fashion. The combination of public and private partners allowed the basic research and hypothesis testing to be quickly accelerated for commercial application by the companies yet maintained that the core framework and academic insights remain in the public domain for continued market disruption, competition, and innovation.

59 BASIC BIOLOGICAL SCIENCES↗

Integration of Waveform Simulation Methods

The generation of synthetic seismograms through simulation is a fundamental tool of seismology required to run quantitative hypothesis tests. A variety of approaches have been developed throughout the seismological community and each has their own specific user interface based on their implementation. This causes a challenge to researchers who will need to learn new interfaces with each new software they wish to use and create substantial challenges when attempting to compare results from different tools. Here we provide a unified interface that facilitates interoperability amongst several simulation tools through a modern containerized Python package. Further, this package includes post-processing analysis modules designed to facilitate end-to-end analysis of synthetic seismograms. In this report we present the conceptual guidance and an example implementation of the new Waveform Simulation Framework.

58 GEOSCIENCES↗

Pore architecture controls on mineral reactivity

Mineral dissolution rates measured in natural environments are much slower than those measured in laboratory settings. This project tested the hypothesis that the way fluid flows through rocks in natural systems creates areas where mineral dissolution is fast and areas where mineral dissolution is slow. This hypothesis was tested with a combination of laboratory experiments and numerical simulation. We demonstrated a separation of fluid flow pathways and rates of mineral dissolution in laboratory experiments for the first time using an experimental approach where we created synthetic rocks that have different ratios of connected and dead-end pathways for fluid flow. In laboratory experiments with higher proportions of dead-end pathways, the mineral dissolution rates were slower. We found that where fluid flows through connected pathways the continuous refreshing of fluid at the mineral surface creates conditions where dissolution is fast. In contrast, where fluid either flows slowly through poorly connected pathways or is stagnant in dead-end pathways, mineral dissolution is slow. The results from this project suggest that the overall slowing of rates of mineral dissolution is important when the proportion of dead-end pathways is greater than ~40%. This project informs our understanding of the way that fluids react with rocks in carbon dioxide sequestration and enhanced geothermal projects where fluids are purposefully injected into rocks for energy applications.

58 GEOSCIENCES↗

Formation of Organic Compounds Through Meteoritic Atmospheric Shock

This document is a Final Technical Report for DoE award DE-SC0023375 “Formation of Organic Compounds Through Meteoritic Atmospheric Shock”. The document includes a summary of topics studied, specific tasks completed, challenges, and results from the project. The main goal of this project was to investigate the production of organic molecules and/or complex inorganic precursor molecules in a plasma environment reminiscent of the environment surrounding meteoroids during atmospheric entries. The specific hypothesis tested in this project was that meteoroid ablation during the entry and the chemical reactions in the meteoroid plasma tail could have produced significant amounts of organics or precursor inorganics in the Early Earth’s atmosphere. Investigation of these processes is essential in understanding the origins of life on Earth and the search for life beyond our planet. This project was focused on a set of experiments conducted at the Utilizing the DIII-D tokamak in San Diego, CA. The experiments aimed to study the interaction of carbonaceous and silica materials (typically found in meteoroids) with mixtures of hot plasma gases (mimicking atmospheric entry conditions. The material samples and gas mixtures were selected to investigate the synthesis of the organic compound urea – a key ingredient in the origin of life – or one of its precursor, the complex inorganic compound ammonia.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Estuarine nutrient pollution impact reduction assessment through euphotic zone avoidance/bypass considerations

The feasibility of reducing nutrient pollution impact by redirecting the effluent to depths below the euphotic zone was investigated for the deep estuarine Puget Sound region of the Salish Sea in the Pacific Northwest of America. The hypothesis tested was that the thickness of the outflow layer in deep estuaries may be greater than the euphotic zone depth, allowing a fraction of nutrients to be exported out passively through the layers immediately below. The euphotic zone depth in Puget Sound varies from 8 to 25 m while the depth of the outflow layer can reach up to ≈ 60 m. Outfall relocation strategies were tested on 99% of the anthropogenic nutrient loads currently delivered to Puget Sound. The impact was quantified using the previously established biophysical Salish Sea Model, using gross primary production and exposure to low dissolved oxygen (DO) levels as the metric (< 2 mg/L for hypoxia and < 5 mg/L for impairment). Eliminating nutrient pollution (above natural) from rivers and wastewater reduced hypoxia exposure by 8.1% and 11.2%, respectively. Relocating the outfalls to deeper waters resulted in improvements, but only in the sill-less sub-basins such as Whidbey, where hypoxia and DO impairment exposure decreased (7.9% and 6.8%, respectively). The presence of multiple sills and circulation cells in Puget Sound resulted in increased exposure and rendered nutrient bypass goals unfeasible as originally envisioned. However, an alternate nutrient export pathway was identified through bottom exchange flow out of Puget Sound via Whidbey Basin and Deception Pass. An unexpected reduction in the exchange outflow magnitude (≈ 4%) due additional (22%) freshwater discharged to the estuary bottom was also noted. The potential loss in circulation strength due to rerouting of natural surface freshwater through submerged deep-water outfalls is identified as a new unforeseen anthropogenic impact.

54 ENVIRONMENTAL SCIENCES↗

SNM Radiation Signature Classification Using Different Semi-Supervised Machine Learning Models

The timely detection of special nuclear material (SNM) transfers between nuclear facilities is an important monitoring objective in nuclear nonproliferation. Persistent monitoring enabled by successful detection and characterization of radiological material movements could greatly enhance the nuclear nonproliferation mission in a range of applications. Supervised machine learning can be used to signal detections when material is present if a model is trained on sufficient volumes of labeled measurements. However, the nuclear monitoring data needed to train robust machine learning models can be costly to label since radiation spectra may require strict scrutiny for characterization. Therefore, this work investigates the application of semi-supervised learning to utilize both labeled and unlabeled data. As a demonstration experiment, radiation measurements from sodium iodide (NaI) detectors are provided by the Multi-Informatics for Nuclear Operating Scenarios (MINOS) venture at Oak Ridge National Laboratory (ORNL) as sample data. Anomalous measurements are identified using a method of statistical hypothesis testing. After background estimation, an energy-dependent spectroscopic analysis is used to characterize an anomaly based on its radiation signatures. In the absence of ground-truth information, a labeling heuristic provides data necessary for training and testing machine learning models. Supervised logistic regression serves as a baseline to compare three semi-supervised machine learning models: co-training, label propagation, and a convolutional neural network (CNN). In each case, the semi-supervised models outperform logistic regression, suggesting that unlabeled data can be valuable when training and demonstrating value in semi-supervised nonproliferation implementations.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

The resolution of point sources of light as analyzed by quantum detection theory

The resolvability of point sources of incoherent light is analyzed by quantum detection theory in terms of two hypothesis-testing problems. In the first, the observer must decide whether there are two sources of equal radiant power at given locations, or whether there is only one source of twice the power located midway between them. In the second problem, either one, but not both, of two point sources is radiating, and the observer must decide which it is. The decisions are based on optimum processing of the electromagnetic field at the aperture of an optical instrument. In both problems the density operators of the field under the two hypotheses do not commute. The error probabilities, determined as functions of the separation of the points and the mean number of received photons, characterize the ultimate resolvability of the sources.

Helstrom, C. W.↗

Detection and Estimation of an Optical Image by Photon-Counting Techniques

Statistical description of a photoelectric detector is given. The photosensitive surface of the detector is divided into many small areas, and the moment generating function of the photo-counting statistic is derived for large time-bandwidth product. The detection of a specified optical image in the presence of the background light by using the hypothesis test is discussed. The ideal detector based on the likelihood ratio from a set of numbers of photoelectrons ejected from many small areas of the photosensitive surface is studied and compared with the threshold detector and a simple detector which is based on the likelihood ratio by counting the total number of photoelectrons from a finite area of the surface. The intensity of the image is assumed to be Gaussian distributed spatially against the uniformly distributed background light. The numerical approximation by the method of steepest descent is used, and the calculations of the reliabilities for the detectors are carried out by a digital computer.

Wang, Lily Lee↗

Resolution of point sources of light as analyzed by quantum detection theory.

The resolvability of point sources of incoherent thermal light is analyzed by quantum detection theory in terms of two hypothesis-testing problems. In the first, the observer must decide whether there are two sources of equal radiant power at given locations, or whether there is only one source of twice the power located midway between them. In the second problem, either one, but not both, of two point sources is radiating, and the observer must decide which it is. The decisions are based on optimum processing of the electromagnetic field at the aperture of an optical instrument. In both problems the density operators of the field under the two hypotheses do not commute. The error probabilities, determined as functions of the separation of the points and the mean number of received photons, characterize the ultimate resolvability of the sources.

Helstrom, C. W.↗

Control of Finite-State, Finite Memory Stochastic Systems

A generalized problem of stochastic control is discussed in which multiple controllers with different data bases are present. The vehicle for the investigation is the finite state, finite memory (FSFM) stochastic control problem. Optimality conditions are obtained by deriving an equivalent deterministic optimal control problem. A FSFM minimum principle is obtained via the equivalent deterministic problem. The minimum principle suggests the development of a numerical optimization algorithm, the min-H algorithm. The relationship between the sufficiency of the minimum principle and the informational properties of the problem are investigated. A problem of hypothesis testing with 1-bit memory is investigated to illustrate the application of control theoretic techniques to information processing problems.

Sandell, Nils R.↗

A self-reorganizing digital flight control system for aircraft

This paper presents a design method for digital self-reorganizing control systems which is optimally tolerant of failures in aircraft sensors. The functions of this system are accomplished with software instead of the popular and costly technique of hardware duplication. The theoretical development, based on M-ary hypothesis testing, results in a bank of M Kalman filters operating in parallel in the failure detection logic. A moving window of the innovations of each Kalman filter drives the detection logic to decide the failure state of the system. The detection logic also selects the optimal state estimate (for control logic) from the bank of Kalman filters. The design process is applied to the design of a self-reorganizing control system for a current configuration of the space shuttle orbiter at Mach 5 and 120,000 feet. The failure detection capabilities of the system are demonstrated using a real-time simulation of the system with noisy sensors.

Montgomery, R. C.↗

Comparison of some biased estimation methods (including ordinary subset regression) in the linear model

Ridge, Marquardt's generalized inverse, shrunken, and principal components estimators are discussed in terms of the objectives of point estimation of parameters, estimation of the predictive regression function, and hypothesis testing. It is found that as the normal equations approach singularity, more consideration must be given to estimable functions of the parameters as opposed to estimation of the full parameter vector; that biased estimators all introduce constraints on the parameter space; that adoption of mean squared error as a criterion of goodness should be independent of the degree of singularity; and that ordinary least-squares subset regression is the best overall method.

Sidik, S. M.↗