Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hypothesis tests”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

From quantitative trait loci towards mechanisms: Linkage Integration Hypothesis Testing (LIgHT) sheds light on the mechanisms of genetically modulated stress tolerance

The goal of this work is to assess the mechanistic bases of natural genetic variations in plant responses of photosynthesis to stress. To achieve this goal, we devised the Linkage Integration Hypothesis Testing (LIgHT) approach, comparing chromosomal locations of quantitative trait loci (QTLs) for multiple phenotypes to distinguish between hypothetical mechanisms. As a use case, we explored genetic variations in photosynthesis-related processes under chilling stress in recombinant inbred lines of cowpea ( Vigna unguiculata L. Walp.). We focused on photosynthesis-related parameters measurable in high throughput and indicative of proposed chilling responses, including the states of PSI and PSII, photoprotective non-photochemical quenching, PSII photodamage, and nyctinastic leaf movements (NLMs). The patterns of QTL linkages indicated that chilling stress tolerance is genetically controlled by avoiding PSII photodamage rather than PSI damage or NLMs. This model was validated in a separate experiment measuring the rates of PSII photodamage and repair. Additional linkages suggest that chilling-induced damage to PSII is controlled by the thylakoid proton motive force and redox state of PSII. This regulation appears to be modulated by thylakoid fatty acid composition, previously associated with the same genetic loci and now supported by broader mechanistic evidence. We propose that the LIgHT approach can be broadly applied to test mechanisms underlying genetic variations.

MultispeQ↗

Hypothesis testing via AI: Generating physically interpretable models of scientific data with machine learning (Full Technical Report)

Deep learning has demonstrated an exceptional ability to solve complex tasks (an engineering success); however, it has done so at the expense of the ability to generate new knowledge (a scientific failure). We propose an alternative framework—entitled Deep Symbolic Regression (DSR)—in which artificial neural networks (NNs) rapidly generate hypotheses about physical relationships among inputs. This framework bypasses the need to interpret an NN altogether, while still leveraging the representational power of deep learning. The resulting models are tractable mathematical expressions, which are inherently and readily human interpretable and can provide insights into underlying physical phenomena. Further, we fold this methodology into the scientific process by allowing the scientist to directly integrate a priori knowledge and beliefs to accelerate learning. We demonstrate this methodology on symbolic regression—the problem of rediscovering underlying expressions describing a dataset—and achieve state-of-the-art performance across a wide variety of symbolic regression problems. Further, we generalize our DSR framework to apply to the more general class of symbolic optimization problems, in which one seeks to optimize a sequence of symbols or “tokens” under a black-box reward function. Examples of other symbolic optimization problems include neural architecture search and computational antibody design. Our generalized tool, Deep Symbolic Optimization (DSO), has been demonstrated on the task of learning symbolic control policies for reinforcement learning environments, and has been adopted as an enabling capability for computational antibody design.

97 MATHEMATICS AND COMPUTING↗

Statistical Significance Testing for Mixed Priors: A Combined Bayesian and Frequentist Analysis

In many hypothesis testing applications, we have mixed priors, with well-motivated informative priors for some parameters but not for others. The Bayesian methodology uses the Bayes factor and is helpful for the informative priors, as it incorporates Occam’s razor via the multiplicity or trials factor in the look-elsewhere effect. However, if the prior is not known completely, the frequentist hypothesis test via the false-positive rate is a better approach, as it is less sensitive to the prior choice. We argue that when only partial prior information is available, it is best to combine the two methodologies by using the Bayes factor as a test statistic in the frequentist analysis. We show that the standard frequentist maximum likelihood-ratio test statistic corresponds to the Bayes factor with a non-informative Jeffrey’s prior. We also show that mixed priors increase the statistical power in frequentist analyses over the maximum likelihood test statistic. We develop an analytic formalism that does not require expensive simulations and generalize Wilks’ theorem beyond its usual regime of validity. In specific limits, the formalism reproduces existing expressions, such as the p-value of linear models and periodograms. We apply the formalism to an example of exoplanet transits, where multiplicity can be more than 10 7 . We show that our analytic expressions reproduce the $p$-values derived from numerical simulations. We offer an interpretation of our formalism based on the statistical mechanics. We introduce the counting of states in a continuous parameter space using the uncertainty volume as the quantum of the state. We show that both the $p$-value and Bayes factor can be expressed as an energy versus entropy competition.

97 MATHEMATICS AND COMPUTING↗

Anomaly Detection in Power System State Estimation: Review and New Directions

Foundational and state-of-the-art anomaly-detection methods through power system state estimation are reviewed. Traditional components for bad data detection, such as chi-square testing, residual-based methods, and hypothesis testing, are discussed to explain the motivations for recent anomaly-detection methods given the increasing complexity of power grids, energy management systems, and cyber-threats. In particular, state estimation anomaly detection based on data-driven quickest-change detection and artificial intelligence are discussed, and directions for research are suggested with particular emphasis on considerations of the future smart grid.

42 ENGINEERING↗

Particle Filter Based Inference Testing

The primary intent of PAR-FIT (Particle Filter based Inference Testing) is to provide hard inductive evidence that a machine learning model is capable and proven for an individual test input. By examining training data used to form the underlying model functional correlation, an estimate of the reliability that a model will make the correct prediction can be made. The Sequential Probability Ratio Test is used to derive a qualitative evaluation for reliability based on hypothesis testing. The PAR-FIT framework achieves this by implementing a particle filter and the sequential probability ratio test algorithms on the machine learning model training data to determine relevancy of new individual test samples to the training dataset. The kernel function evaluates the local proximity and density of training data used to derive a prediction outcome. Particles are used to probabilistically determine which training data to evaluate for proximity. For test samples that are within a close proximity to and surrounded by multiple training data points, the evaluated reliability of the prediction is high. For test samples that are anomalies not represented by the training dataset, in low density data clusters, or are far from existing data points, the evaluated reliability is low as insufficient training evidence exists to suggest the model is capable of making the correct prediction. Sequential Probability Ratio Test is further used to determine when a hypothesis on whether a signal can be rejected or accepted for use. The ratio test collects sequence information from the particle filter to test whether the signal is anomalous or normal via hypothesis testing of the underlying distributions.

Chen, Edward [Idaho National Laboratory (INL), Ida↗

A novel framework for increasing research transparency: Exploring the connection between diversity and innovation

A split sample/dual method research protocol is demonstrated to increase transparency while reducing the probability of false discovery. We apply the protocol to examine whether diversity in ownership teams increases or decreases the likelihood of a firm reporting a novel innovation using data from the 2018 United States Census Bureau’s Annual Business Survey. Transparency is increased in three ways: 1) all specification testing and identifying potentially productive models is done in an exploratory subsample that 2) preserves the validity of hypothesis test statistics fromde novoestimation in the holdout confirmatory sample with 3) all findings publicly documented in an earlier registered report and in this journal publication. Bayesian estimation procedures that leverage information from the exploratory stage included in the confirmatory stage estimation replace traditional frequentist null hypothesis significance testing. In addition to increasing statistical power by using information from the full sample, Bayesian methods directly estimate a probability distribution for the magnitude of an effect, allowing much richer inference. Estimated magnitudes of diversity along academic discipline, race, ethnicity, and foreign-born status dimensions are positively associated with innovation. A maximally diverse ownership team on these dimensions would be roughly six times more likely to report new-to-market innovation than a homophilic team.

Science & Technology - Other Topics↗

Separating Sensor Anomalies From Process Anomalies in Data-Driven Anomaly Detection

Data-driven anomaly detection over time series data is studied from the perspective of separating data anomalies—corresponding to sensor failures—from process anomalies—that arise from equipment or operational failures. Herein, a semi-supervised approach is proposed that utilizes two predictive models trained on non-anomalous data using two different sensor groups as inputs, and a nested hypothesis test to reliably classify data or process anomalies. Conditions are derived on choice of sensor groups to guarantee reliable detection, and a case study is presented to demonstrate the proposed classification approach.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Walking the ‘design–build–test–learn’ cycle: flux analysis and genetic engineering reveal the pliability of plant central metabolism

Oilseeds are of great economic importance for food and animal feed and their contribution to renewable energy production. Soybean seeds (Glycine max (L.) Merr.) contain c. 40% protein, 20% oil, and 30% carbohydrate (Song et al., 2023). Due to the massive scale of soybean production worldwide, even small improvements in seed protein and oil content make economic sense (Song et al., 2023). Successful manipulation of seed composition largely depends on a thorough understanding of the processes and pathways involved in the biosynthesis of fatty acids and amino acids, which are the building blocks of lipids and proteins. Rational engineering of the synthesis of storage reserves, that is, the rerouting of metabolic flux in central metabolism, is difficult to accomplish due to the complexity of the central metabolic network, the intricate regulation of its enzymes at multiple levels, and the often-unpredictable effects of genetic manipulation (Sweetlove et al., 2017). Therefore, the advancement of our understanding of central metabolism and its control of carbon partitioning requires following an iterative ‘design–build–test–learn’ (DBTL) cycle (Lin & Eudes, 2020) where metabolic flux analysis and hypothesis testing by transgenic approaches are important components. Previous metabolic studies on soybeans using isotopic tracers and metabolic flux analysis have provided insight into how lipid and protein biosynthesis occurs simultaneously during seed development (Allen et al., 2009; Allen & Young, 2013; Kambhampati et al., 2021). In an article published in this issue of New Phytologist, Morley et al. (2023; 1834–1851) put the insights they have gained into the delivery of metabolic precursors and energy cofactors to oil synthesis to the test and arrive at a successful metabolic engineering design. They show that an increase in seed oil content in soybeans can be achieved by overexpression of malic enzyme (ME) during seed development. Malic enzyme refers to a class of decarboxylating malate dehydrogenase enzymes that oxidize malate with NAD + or NADP + as redox cofactor while generating pyruvate and CO 2 . Like higher plants in general, soybean has distinct NADH- or NADPH-producing ME isoforms localized to the cytosol, plastid, or mitochondria (Gerrard Wheeler et al., 2016). As Morley et al. show, an increase in seed oil can be achieved in particular when a NADP+-dependent enzyme isoform (EC 1.1.1.40) is overexpressed in the plastid. Given the complex compartmentalization of pyruvate, malate, and redox metabolism (Fig. 1), increased oil production appears to depend on additional pyruvate and reducing equivalents being produced in the same compartment where de novo fatty acid biosynthesis occurs: the plastid.

59 BASIC BIOLOGICAL SCIENCES↗

Computational Imaging for Intelligence in Highly Scattering Aerosols (Final Report)

Natural and man-made degraded visual environments pose major threats to national security. The random scattering and absorption of light by tiny particles suspended in the air reduces situational awareness and causes unacceptable down-time for critical systems and operations. To improve the situation, we have developed several approaches to interpret the information contained within scattered light to enhance sensing and imaging in scattering media. These approaches were tested at the Sandia National Laboratory Fog Chamber facility and with tabletop fog chambers. Computationally efficient light transport models were developed and leveraged for computational sensing. The models are based on a weak angular dependence approximation to the Boltzmann or radiative transfer equation that appears to be applicable in both the moderate and highly scattering regimes. After the new model was experimentally validated, statistical approaches for detection, localization, and imaging of objects hidden in fog were developed and demonstrated. A binary hypothesis test and the Neyman-Pearson lemma provided the highest theoretically possible probability of detection for a specified false alarm rate and signal-to-noise ratio. Maximum likelihood estimation allowed estimation of the fog optical properties as well as the position, size, and reflection coefficient of an object in fog. A computational dehazing approach was implemented to reduce the effects of scatter on images, making object features more readily discernible. We have developed, characterized, and deployed a new Tabletop Fog Chamber capable of repeatably generating multiple unique fog-analogues for optical testing in degraded visual environments. We characterized this chamber using both optical and microphysical techniques. In doing so we have explored the ability of droplet nucleation theory to describe the aerosols generated within the chamber, as well as Mie scattering theory to describe the attenuation of light by said aerosols, and correlated the aerosol microphysics to optical properties such as transmission and meteorological optical range (MOR). This chamber has proved highly valuable and has supported multiple efforts inclusive to and exclusive of this LDRD project to test optics in degraded visual environments. Circularly polarized light has been found to maintain its polarization state better than linearly polarized light when propagating through fog. This was demonstrated experimentally in both the visible and short-wave infrared (SWIR) by imaging targets made of different commercially available retroreflective films. It was found that active circularly polarized imaging can increase contrast and range compared to linearly polarized imaging. We have completed an initial investigation of the capability for machine learning methods to reduce the effects of light scattering when imaging through fog. Previously acquired experimental long-wave images were used to train an autoencoder denoising architecture. Overfitting was found to be a problem because of lack of variability in the object type in this data set. The lessons learned were used to collect a well labeled dataset with much more variability using the Tabletop Fog Chamber that will be available for future studies. We have developed several new sensing methods using speckle intensity correlations. First, the ability to image moving objects in fog was shown, establishing that our unique speckle imaging method can be implemented in dynamic scattering media. Second, the speckle decorrelation over time was found to be sensitive to fog composition, implying extensions to fog characterization. Third, the ability to distinguish macroscopically identical objects on a far-subwavelength scale was demonstrated, suggesting numerous applications ranging from nanoscale defect detection to security. Fourth, we have shown the capability to simultaneously image and localize hidden objects, allowing the speckle imaging method to be effective without prior object positional information. Finally, an interferometric effect was presented that illustrates a new approach for analyzing speckle intensity correlations that may lead to more effective ways to localize and image moving objects. All of these results represent significant developments that challenge the limits of the application of speckle imaging and open important application spaces. A theory was developed and simulations were performed to assess the potential transverse resolution benefit of relative motion in structured illumination for radar systems. Results for a simplified radar system model indicate that significant resolution benefits are possible using data from scanning a structured beam over the target, with the use of appropriate signal processing.

58 GEOSCIENCES↗

Demonstrating Hierarchical System Development With the Common Community Physics Package Single‐Column Model: A Case Study Over the Southern Great Plains

This study demonstrates a specific application of the hierarchical system development (HSD) approach to investigate, analyze, and attribute model issues within the Unified Forecast System (UFS), with a focus on process isolation. By evaluating a non‐precipitating, shallow cumulus case at the Atmospheric Radiation Measurement Southern Great Plains site in the UFS global forecast against the observation, the investigation identifies a warmer and deeper daytime convective planetary boundary layer (PBL) and misrepresented nocturnal PBL transition. Hypothesis testing, which employs the Common Community Physics Package (CCPP) single‐column model (SCM) and uses the same physics as the UFS global model, confirms that these issues are attributed to the model physics and initialization. Specifically, misrepresented PBL processes are linked to problematic surface condition and a lack of cloud formation, which may stem from deficiencies in PBL and cloud microphysics parameterizations and their interactions. The UFS initial condition contributes to an earlier, excessively collapsed daytime convective boundary layer and a lack of decoupling between the stable boundary layer and residual layer late in the afternoon. This work introduces an avenue for the community to engage with the application of HSD, along with the CCPP and CCPP SCM, to understand the interplay of model physics, disentangle the roles of model components, as well as facilitate model and forecast improvement.

54 ENVIRONMENTAL SCIENCES↗

Learning likelihood ratios with neural network classifiers

The likelihood ratio is a crucial quantity for statistical inference in science that enables hypothesis testing, construction of confidence intervals, reweighting of distributions, and more. Many modern scientific applications, however, make use of data- or simulation-driven models for which computing the likelihood ratio can be very difficult or even impossible. By applying the so-called “likelihood ratio trick,” approximations of the likelihood ratio may be computed using clever parametrizations of neural network-based classifiers. A number of different neural network setups can be defined to satisfy this procedure, each with varying performance in approximating the likelihood ratio when using finite training data. We present a series of empirical studies detailing the performance of several common loss functionals and parametrizations of the classifier output in approximating the likelihood ratio of two univariate and multivariate Gaussian distributions as well as simulated high-energy particle physics datasets.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Space-time generalization of mutual information

The mutual information characterizes correlations between spatially separated regions of a system. Yet, in experiments we often measure dynamical correlations, which involve probing operators that are also separated in time. Here, we introduce a space-time generalization of mutual information which, by construction, satisfies several natural properties of the mutual information and at the same time characterizes correlations across subsystems that are separated in time. In particular, this quantity, that we call the space-time mutual information, bounds all dynamical correlations. We construct this quantity based on the idea of the quantum hypothesis testing. As a by-product, our definition provides a transparent interpretation in terms of an experimentally accessible setup. We draw connections with other notions in quantum information theory, such as quantum channel discrimination. Finally, we study the behavior of the space-time mutual information in several settings and contrast its long-time behavior in many-body localizing and thermalizing systems.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

High-dimensional isotomics, part 2: Observations of over 100 constraints on methionine's isotome

The abundances of different isotopic forms of a compound, or isotopologues, will vary based on its physical and chemical history. The number of isotopologues increases combinatorically with the size of a molecule, and even small molecules such as amino acids have thousands of potentially observable isotopic variants. However, due to the analytical challenges of separating and observing isotopologues, only a few dimensions of isotopic diversity are routinely measured. Overcoming these challenges requires both an experimental method to observe many isotopic properties and a theoretical framework for interpreting these experiments. In Part 1, we presented such a theoretical framework; here, we demonstrate an experimental method, which we apply to methionine. Our approach uses a Q Exactive HF Orbitrap to perform several “M + N experiments”, where a sample is ionized, a subset of its isotopologues with cardinal mass N daltons greater than the unsubstituted isotopologue is selected and fragmented, and the proportions of all detectable isotopic forms of those fragment ions are quantified. We perform M + 1, M + 2, M + 3, and M + 4 experiments of a methionine sample and standard where the sample has a 100 ‰ enrichment of 13C at the methyl carbon relative to the natural 13C abundance at that position in the standard, and is otherwise identical to the standard. We observe isotopic forms of 8 fragment ion species for each version of the M + N experiment. With the assistance of a forward model of expected mass spectra, we identify isotopic peaks for each fragment ion based on observed mass and abundance, screen these for data quality, and quantify abundances for 146 unique isotopic peaks at precisions of ≈ 0.3–3 ‰. We present our direct observations and use them to reconstruct the concentrations of 19 individual singly, doubly, and triply-substituted isotopologues; doing so gives fewer constraints and broader error bars than working with the direct observations, but may be more interpretable for some applications. We also examine possibilities for measuring additional peaks, which are primarily limited by the detection limit of the Orbitrap-IRMS method. We then suggest some possible uses of our direct measurements for chemical forensics and hypothesis testing. Furthermore, our results demonstrate the diversity of isotopic constraints currently observable and interpretable for organic molecules.

58 GEOSCIENCES↗

Hydrogen underground storage for grid electricity storage: An optimization study on techno-economic analysis

Here, this study performs a techno-economic analysis of hydrogen underground storage systems for grid electricity storage, evaluating their economic viability at the plant scale using dynamic optimization. It explores the feasibility of various system configurations and revenue models in the context of volatile electricity prices and the necessity for multiple revenue streams. The hypothesis tested is that large-scale hydrogen storage, despite its low round-trip efficiency, can be economically viable with the right mix of revenue streams. This study uses scenario-based analysis to assess the impacts of different system configurations, including engaging in time-shifting arbitrage, ancillary service markets and blending hydrogen with natural gas. Results indicate potential annual net cash flows of up to $\$$1.5 million from ancillary services integration and $\$$5.2 million from natural gas blending, contingent on specific system sizes. The study concludes that hydrogen underground storage for grid electricity storage can be profitable, and emphasizes that proper system design and precise electricity price forecasting are crucial for optimizing system performance and economic returns. This research sets the stage for further investigations into the scalability of hydrogen storage systems and their broader implications for grid electricity storage and energy market dynamics.

25 ENERGY STORAGE↗

Dynamic and single cell characterization of a CRISPR-interference toolset in Pseudomonas putida KT2440 for β-ketoadipate production from p -coumarate

We report Pseudomonas putida KT2440 is a well-studied bacterium for the conversion of lignin-derived aromatic compounds to bioproducts. The development of advanced genetic tools in P. putida has reduced the turnaround time for hypothesis testing and enabled the construction of strains capable of producing various products of interest. Here, we evaluate an inducible CRISPR-interference (CRISPRi) toolset on fluorescent, essential, and metabolic targets. Nuclease-deficient Cas9 (dCas9) expressed with the arabinose (8K)-inducible promoter was shown to be tightly regulated across various media conditions and when targeting essential genes. In addition to bulk growth data, single cell time lapse microscopy was conducted, which revealed intrinsic heterogeneity in knockdown rate within an isoclonal population. The dynamics of knockdown were studied across genomic targets in exponentially-growing cells, revealing a universal 1.75 ± 0.38 hour quiescent phase after induction where 1.5 ± 0.35 doublings occur before a phenotypic response is observed. To demonstrate application of this CRISPRi toolset, β-ketoadipate, a monomer for performance-advantaged nylon, was produced at a 4.39 ± 0.5 g/L and yield of 0.76 ± 0.10 mol/mol from p-coumarate, a hydroxycinnamic acid that can be derived from grasses. These cultivation metrics were achieved by using the higher strength IPTG (1K)-inducible promoter to knockdown the pcaIJ operon in the βKA pathway during early exponential phase. This allowed the majority of the carbon to be shunted into the desired product while eliminating the need for a supplemental carbon and energy source to support growth and maintenance.

59 BASIC BIOLOGICAL SCIENCES↗

Multivariable degradation modeling and life prediction using multivariate fractional Brownian motion

In system prognostics and health management, multivariable degradation models have been widely developed to predict the life of complex systems using degradation data of multiple Performance Characteristics (PCs). Recent studies have detected a Long-Term Memory (LTM) effect among the degradation process of various PCs, implying a strong coupling phenomenon between the future degradation behavior and historical degradation trajectory. Although the LTM has been widely integrated into single-PC-based degradation modeling, it has not been considered in multi-PC-based scenarios. To capture LTM among multiple PCs, this article proposes a novel LTM-integrated Multivariate Degradation Model (MDM) for system life prediction based on multivariate fractional Brownian motion, which simultaneously incorporates the cross-correlation among different PCs. To estimate parameters of the LTM-integrated MDM, a maximum likelihood method is developed. Here, two likelihood-ratio hypothesis tests are developed to test the existence of the overall and individual LTM effect among multiple PCs. Both simulation studies and physical experiments on the performance degradation of solar energy conversion and storage devices are conducted to validate the proposed model. Results reveal that the proposed LTM-integrated MDM significantly outperforms existing MDMs in life prediction, while the lifetime uncertainty is heavily underestimated by those traditional approaches that neglect the LTM.

42 ENGINEERING↗

A Mass‐Conserving‐Perceptron for Machine‐Learning‐Based Modeling of Geoscientific Systems

Although decades of effort have been devoted to building Physical-Conceptual (PC) models for predicting the time-series evolution of geoscientific systems, recent work shows that Machine Learning (ML) based Gated Recurrent Neural Network technology can be used to develop models that are much more accurate. However, the difficulty of extracting physical understanding from ML-based models complicates their utility for enhancing scientific knowledge regarding system structure and function. Here, we propose a physically interpretable Mass-Conserving-Perceptron (MCP) as a way to bridge the gap between PC-based and ML-based modeling approaches. The MCP exploits the inherent isomorphism between the directed graph structures underlying both PC models and GRNNs to explicitly represent the mass-conserving nature of physical processes while enabling the functional nature of such processes to be directly learned (in an interpretable manner) from available data using off-the-shelf ML technology. As a proof of concept, we investigate the functional expressivity (capacity) of the MCP, explore its ability to parsimoniously represent the rainfall-runoff (RR) dynamics of the Leaf River Basin, and demonstrate its utility for scientific hypothesis testing. To conclude, we discuss extensions of the concept to enable ML-based physical-conceptual representation of the coupled nature of mass-energy-information flows through geoscientific systems.

58 GEOSCIENCES↗