Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Statistical accuracy”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Mitigating the noise of DESI mocks using analytic control variates

In order to address fundamental questions related to the expansion history of the Universe and its primordial nature with the next generation of galaxy experiments, we need to model reliably large-scale structure observables such as the correlation function and the power spectrum. Cosmological N-body simulations provide a reference through which we can test our models, but their output suffers from sample variance on large scales. Fortunately, this is the regime where accurate analytic approximations exist. To reduce the variance, which is key to making optimal use of these simulations, we can leverage the accuracy and precision of such analytic descriptions using Control Variates (CV). The power of control variates stems from utilizing inexpensive but highly correlated surrogates of the statistics one wishes to measure. The stronger the correlation between the surrogate and the statistic of interest, the larger the variance reduction delivered by the method. We apply two control variate formulations to mock catalogs generated in anticipation of upcoming data from the Dark Energy Spectroscopic Instrument (DESI) to test the robustness of its analysis pipeline. Our CV-reduced measurements offer a factor of 5-10 improvement in the measurement error compared with the raw measurements. We explore the relevant properties of the galaxy samples that dictate this reduction and comment on the improvements we find on some of the derived quantities relevant to Baryon Acoustic Oscillation (BAO) analysis.

79 ASTRONOMY AND ASTROPHYSICS↗

Analysis of Warped April Tag Impacts on Detection and Pose Estimation

This report evaluates the impact of geometric deformation on an April Tag, particularly when warped due to attachment on a curved surface, on its detectability and pose estimation performance. A comparative analysis was conducted using a flat April Tag as a control under identical experimental conditions, which involved recording video sequences with varying viewing angles. For detectability, the warped tag exhibited consistent detection failures at viewing angles beyond 40° and complete failures beyond 60°, whereas the flat tag maintained reliable detection across all angles. For pose estimation, measured by pose jitter (variation in rotation and translation), differences between the warped and flat tags were minimal and statistically insignificant, indicating robust performance even for the deformed tag. These findings suggest that while geometric warping reduces an April Tag’s detectability, its pose estimation accuracy remains relatively unaffected under the tested conditions.

42 ENGINEERING↗

Feeling the Strain: Quantifying Ligand Deformation in Photosynthesis

Structural distortion of protein-bound ligands can play a critical role in enzyme function by tuning the electronic and chemical properties of the ligand molecule. However, quantifying these effects is difficult due to the limited resolution of protein structures and the difficulty of generating accurate structural restrains for non-protein ligands. Here, we seek to quantify these effects through a statistical analysis of ligand distortion in Chlorophyll (Chl) proteins (CP), where ring deformation is thought to play a role in energy and electron transfer. To assess the accuracy of ring-deformation estimates from available structural data, we take advantage of the C 2 symmetry of Photosystem II (PSII), comparing ring-deformation estimates for equivalent sites both within and between 113 distinct X-ray and Cryogenic electron microscopy (CryoEM) PSII structures. Significantly, we find that several deformation modes exhibit considerable variability in predictions, even for equivalent monomers, down to 2 Å resolution, to an extent that probably prevents their utilization in optical calculations. We further find that refinement restrains play a critical role in determining deformation values to resolution as low as 2 Å. However, for those modes that are well-resolved in the structural data, ring deformation in PSII is strongly conserved across all species tested, from cyanobacteria to algae. Furthermore, these results highlight both the opportunities and limitations inherent in the structure-based analysis of the bioenergetic and optical properties of CPs and other protein-ligand complexes.

14 SOLAR ENERGY↗

Combining physics-based and data-driven models for quantitatively accurate plasma profile prediction that extrapolates well; with application to DIII-D, AUG, and ITER tokamaks

For design, scenario planning, and control, ITER and all other envisioned tokamaks rely on a variety of statistical and physics-based models to extrapolate to unseen regimes; most notably from low plasma current to high. A 'meta-learning' methodology for combining the accuracy of data-driven models with the generalizability of physics-based models is described and tested, yielding a 5–10 percent improvement in performance beyond either alone for the task of extrapolating time-dependent plasma profile prediction from low- to high- plasma current DIII-D tokamak discharges. Meanwhile, it is shown that both machine learning models extrapolated far-distribution and state-of-the-art 'physics-based' profile predictors fare worse than merely assuming plasma profiles do not change from their initial values. Finally, a variety of other mechanisms for helping data-driven models generalize—transfer learning, adding contextual information from physics simulators, and adding data from the ASDEX Upgrade tokamak—are attempted for similar extrapolation tasks but, in the methodology used in this paper, yield no significant improvement beyond simple data-driven models. Results are summarized in figures 15 and 16.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Climate-invariant machine learning

Projecting climate change is a generalization problem: We extrapolate the recent past using physical models across past, present, and future climates. Current climate models require representations of processes that occur at scales smaller than model grid size, which have been the main source of model projection uncertainty. Recent machine learning (ML) algorithms hold promise to improve such process representations but tend to extrapolate poorly to climate regimes that they were not trained on. To get the best of the physical and statistical worlds, we propose a framework, termed “climate-invariant” ML, incorporating knowledge of climate processes into ML algorithms, and show that it can maintain high offline accuracy across a wide range of climate conditions and configurations in three distinct atmospheric models. Our results suggest that explicitly incorporating physical knowledge into data-driven models of Earth system processes can improve their consistency, data efficiency, and generalizability across climate regimes.

54 ENVIRONMENTAL SCIENCES↗

Detectability of Varied Hybridization Scenarios Using Genome-Scale Hybrid Detection Methods

Hybridization events complicate the accurate reconstruction of phylogenies, as they lead to patterns of genetic heritability that are unexpected under traditional, bifurcating models of species trees. This phenomenon has led to the development of methods to infer these varied hybridization events, both methods that reconstruct networks directly, as well as summary methods that predict individual hybridization events from a subset of taxa. However, a lack of empirical comparisons between methods – especially those pertaining to large networks with varied hybridization scenarios – hinders their practical use. Here, we provide a comprehensive review of popular summary methods: TICR, MSCquartets, HyDe, Patterson’s D-Statistic (ABBA-BABA), D3, and Dp. TICR and MSCquartets are based on quartet concordance factors gathered from gene tree topologies and HyDe, Patterson’s D-Statistic, D3, and Dp use site pattern frequencies to identify hybridization events between sets of three taxa. We then use simulated data to address questions of method accuracy and ideal use scenarios by testing methods against complex networks which depict gene flow events that differ in depth (timing), quantity (single vs. multiple, overlapping hybridizations), and rate of gene flow (γ). We find that deeper or multiple hybridization events may introduce noise and weaken the signal of hybridization, leading to higher relative false negative rates across all methods. Despite some forms of hybridization eluding quartet-based detection methods, MSCquartets displays high precision in most scenarios. While HyDe results in high false negative rates when tested on hybridizations involving extinct or unsampled ghost lineages, HyDe is the only method able to identify the direction of hybridization, distinguishing the source parental lineages from recipient hybrid lineages. Lastly, we test the methods on a dataset of ultraconserved elements from the bee subfamily Nomiinae, finding possible hybridization events between clades which correspond to regions of poor support in the species tree estimated in a previous study.

Bjorner, Marianne B.↗

Lithium-ion battery physics and statistics-based state of health model

A pseudo-2d model using COMSOL Multiphysics® software is developed to simulate performance and performance degradation of Li-ion batteries consisting of layered and olivine cathodes with graphite anode when subjected to peak shaving grid service. Multiple degradation pathways are considered, including solid electrolyte interphase (SEI) formation and breakdown at the anode, cathode dissolution and its synergistic effect on SEI formation at the anode. The model is validated by simulating commercial cylindrical cell performance. A global model is developed to simulate performance across all chemistries, along with individual chemistry models using global model parameters as initial values. There is good agreement between these models for various optimization parameters such as SEI equilibrium potential, cathode dissolution exchange current density, solvent diffusivity in the SEI and SEI ionic conductivity. To circumvent time constraints related to the COMSOL model, a 0d global model is developed which fits data well and provides more clarity on differences in cathode dissolution exchange current density. Again, good agreement for various optimization parameters is obtained among the COMSOL global & individual chemistry models and the 0-d model. The lessons learned from the physics-based model is used to develop a top down statistics-based model using current, voltage and anode volumetric change per mole lithium intercalated, along with their interactions as degradation predictors. This model predicts out of sample degradation for multiple grid services and electric vehicle drive cycle with high accuracy and provides the pathway to develop an efficient battery management system combining machine learning and findings from physics-based computationally intensive algorithms.

Crawford, Aladsair J.↗

Identifying and tracking bubbles and drops in simulations: A toolbox for obtaining sizes, lineages, and breakup and coalescence statistics

Knowledge of bubble and drop size distributions in two-phase flows is important for characterizing a wide range of phenomena, including combustor ignition, sonar communication, and cloud formation. The physical mechanisms driving the background flow also drive the time evolution of these distributions. Accurate and robust identification and tracking algorithms for the dispersed phase are necessary to reliably measure this evolution and thereby quantify the underlying mechanisms in interface-resolving flow simulations. The identification of individual bubbles and drops traditionally relies on an algorithm used to identify connected regions. This traditional algorithm can be sensitive to the presence of spurious structures. A cost-effective refinement is proposed to maximize volume accuracy while minimizing the identification of spurious bubbles and drops. An accurate identification scheme is crucial for distinguishing bubble and drop pairs with large size ratios. The identified bubbles and drops need to be tracked in time to obtain breakup and coalescence statistics that characterize the evolution of the size distribution, including breakup and coalescence frequencies, and the probability distributions of parent and child bubble and drop sizes. An algorithm based on mass conservation is proposed to construct bubble and drop lineages using simulation snapshots that are not necessarily from consecutive time steps. These lineages are then used to detect breakup and coalescence events, and obtain the desired statistics. Accurate identification of large-size-ratio bubble and drop pairs enables accurate detection of breakup and coalescence events over a large size range. Accurate detection of successive breakup and coalescence events requires that the snapshot interval be an order of magnitude smaller than the characteristic breakup and coalescence times to capture these successive events while minimizing the identification of repeated confounding events. Together, these algorithms serve as a toolbox for detailed analysis of two-phase simulations, and enable insights into the mechanisms behind bubble and drop formation and evolution in flows of practical importance.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

An evaluation of multi-fidelity methods for quantifying uncertainty in projections of ice-sheet mass change

Abstract. This study investigated the computational benefits of using multi-fidelity statistical estimation (MFSE) algorithms to quantify uncertainty in the mass change of Humboldt Glacier, Greenland, between 2007 and 2100 using a single climate change scenario. The goal of this study was to determine whether MFSE can use multiple models of varying cost and accuracy to reduce the computational cost of estimating the mean and variance of the projected mass change of a glacier. The problem size and complexity were chosen to reflect the challenges posed by future continental-scale studies while still facilitating a computationally feasible investigation of MFSE methods. When quantifying uncertainty introduced by a high-dimensional parameterization of the basal friction field, MFSE was able to reduce the mean-squared error in the estimates of the statistics by well over an order of magnitude when compared to a single-fidelity approach that only used the highest-fidelity model. This significant reduction in computational cost was achieved despite the low-fidelity models used being incapable of capturing the local features of the ice-flow fields predicted by the high-fidelity model. The MFSE algorithms were able to effectively leverage the high correlation between each model's predictions of mass change, which all responded similarly to perturbations in the model inputs. Consequently, our results suggest that MFSE could be highly useful for reducing the cost of computing continental-scale probabilistic projections of sea-level rise due to ice-sheet mass change.

54 ENVIRONMENTAL SCIENCES↗

Constraining gravity with a new precision 𝐸 𝐺 estimator using Planck + SDSS BOSS data

The 𝐸 𝐺 statistic is a discriminating probe of gravity developed to test the prediction of general relativity (GR) for the relation between gravitational potential and clustering on the largest scales in the observable Universe. We present a novel high-precision estimator for the 𝐸 𝐺 statistic using CMB lensing and galaxy clustering correlations that carefully matches the effective redshifts across the different measurement components to minimize corrections. A suite of detailed tests is performed to characterize the estimator’s accuracy, its sensitivity to assumptions and analysis choices, and the non-Gaussianity of the estimator’s uncertainty is characterized. After finalization of the estimator, it is applied to Planck CMB lensing and SDSS CMASS and LOWZ galaxy data. We report the first harmonic space measurement of 𝐸 𝐺 using the LOWZ sample and CMB lensing and also updated constraints using the final CMASS sample and the latest Planck CMB lensing map. We find $\hat{𝐸}$$^{Planck+CMASS}_{𝐺}$ = 0.3⁢6$^{+0.06}_{−0.05}$⁢(68.27%) and $\hat{𝐸}$$^{Planck+LOWZ}_{𝐺}$ = 0.4⁢0$^{+0.11}_{−0.09}$⁢(68.27%), with additional subdominant systematic error budget estimates of 2% and 3%, respectively. Using Ω m,0 constraints from Planck and SDSS BAO observations, Λ⁢CDM-GR predicts 𝐸$^{GR}_ {𝐺}$⁡(𝑧 =0.555) = 0.401 ± 0.005 and 𝐸$^{GR}_{𝐺}$⁡(𝑧 =0.316) = 0.452 ± 0.005 at the effective redshifts of the CMASS and LOWZ based measurements. We report the measurement to be in good statistical agreement with the Λ⁢CDM-GR prediction and report that the measurement is also consistent with the more general GR prediction of scale independence for 𝐸 𝐺 . Furthermore, this work provides a carefully constructed and calibrated statistic with which 𝐸 𝐺 measurements can be confidently and accurately obtained with upcoming survey data.

79 ASTRONOMY AND ASTROPHYSICS↗

Efficient screening of rare large pit anomalies on polished surfaces using a minimalist sampling scheme

Lawrence Livermore National Laboratory (LLNL) has made significant strides in generating clean energy through its inertial confinement fusion (ICF) experiments. These experiments rely on high-density carbon (HDC) coated shells to encapsulate the fusion fuel. The success of these experiments is heavily dependent on the surface quality of these shells, as even minor imperfections, such as deep pits, can negatively impact fusion yield. Ensuring the required smoothness involves an extensive surface-finishing process that spans approximately 20 stages, making it both time-intensive and resource-demanding. A critical challenge in this process is the need for high-resolution scans to detect rare deep pits, which can be costly and impractical if performed on every shell. This highlights the necessity of developing more efficient scanning methods to optimize time and cost without compromising accuracy. To address these challenges, we introduce a novel approach that employs the multivariate Dvoretzky–Kiefer–Wolfowitz (DKW) inequality to provide a probabilistic upper bound on the error in estimating pit distribution characteristics via a Kernel Density Estimator (KDE). This error bound enables efficient and reliable estimation of pit distribution characteristics at a specified statistical confidence level using a minimal number of surface scans. The integrated DKW-KDE approach was validated through surface-finishing experiments across two batches of HDC-coated shells, demonstrating consistent and robust performance across multiple stages of the surface-finishing experiments. The validation studies suggest that the integrated DKW-KDE approach achieves comparable accuracy in estimating the risk of deleterious large pits with six scans, thus conserving time and resources. Further evaluations show that performance remains consistent across batches and over multiple polishing stages. In conclusion, based on these findings, one can leverage the minimal-scan insights to strategically improve the bottleneck inspection process, thus enhancing the productivity and quality of shell polishing and similar challenging manufacturing processes.

Inertial confinement fusion↗

Towards testing the theory of gravity with DESI: summary statistics, model predictions and future simulation requirements

Shortly after its discovery, General Relativity (GR) was applied to predict the behavior of our Universe on the largest scales, and later became the foundation of modern cosmology. Its validity has been verified on a range of scales and environments from the Solar system to merging black holes. However, experimental confirmations of GR on cosmological scales have so far lacked the accuracy one would hope for — its applications on those scales being largely based on extrapolation and its validity there sometimes questioned in the shadow of the discovery of the unexpected cosmic acceleration. Future astronomical instruments surveying the distribution and evolution of galaxies over substantial portions of the observable Universe, such as the Dark Energy Spectroscopic Instrument (DESI), will be able to measure the fingerprints of gravity and their statistical power will allow strong constraints on alternatives to GR. In this paper, based on a set of N-body simulations and mock galaxy catalogs, we study the predictions of a number of traditional and novel summary statistics beyond linear redshift distortions in two well-studied modified gravity models — chameleon f(R) gravity and a braneworld model — and the potential of testing these deviations from GR using DESI. These summary statistics employ a wide array of statistical properties of the galaxy and the underlying dark matter field, including two-point and higher-order statistics, environmental dependence, redshift space distortions and weak lensing. We find that they hold promising power for testing GR to unprecedented precision. The major future challenge is to make realistic, simulation-based mock galaxy catalogs for both GR and alternative models to fully exploit the statistic power of the DESI survey (by matching the volumes and galaxy number densities of the mocks to those in the real survey) and to better understand the impact of key systematic effects. Using these, we identify future simulation and analysis needs for gravity tests using DESI.

79 ASTRONOMY AND ASTROPHYSICS↗

Predictive Modeling of NOx Emissions from Lean Direct Injection of Hydrogen and Hydrogen/Natural Gas Blends Using Flame Imaging and Machine Learning

This research paper explores the use of machine learning to relate images of flame structure and luminosity to measured NOx emissions. Images of reactions produced by 16 aero-engine derived injectors for a ground-based turbine operated on a range of fuel compositions, air pressure drops, preheat temperatures and adiabatic flame temperatures were captured and postprocessed. The experimental investigations were conducted under atmospheric conditions, capturing CO, NO and NOx emissions data and OH* chemiluminescence images from 27 test conditions. The injector geometry and test conditions were based on a statistically designed test plan. These results were first analyzed using the traditional analysis approach of analysis of variance (ANOVA). The statistically based test plan yielded 432 data points, leading to a correlation for NOx emissions as a function of injector geometry, test conditions and imaging responses, with 70.2% accuracy. As an alternative approach to predicting emissions using imaging diagnostics as well as injector geometry and test conditions, a random forest machine learning algorithm was also applied to the data and was able to achieve an accuracy of 82.6%. This study offers insights into the factors influencing emissions in ground-based turbines while emphasizing the potential of machine learning algorithms in constructing predictive models for complex systems.

08 HYDROGEN↗

Stopping criteria for ending autonomous, single detector radiological source searches

While the localization of radiological sources has traditionally been handled with statistical algorithms, such a task can be augmented with advanced machine learning methodologies. The combination of deep and reinforcement learning has provided learning-based navigation to autonomous, single-detector, mobile systems. However, these approaches lacked the capacity to terminate a surveying/search task without outside influence of an operator or perfect knowledge of source location (defeating the purpose of such a system). Two stopping criteria are investigated in this work for a machine learning navigated system: one based upon Bayesian and maximum likelihood estimation (MLE) strategies commonly used in source localization, and a second providing the navigational machine learning network with a “stop search” action. A convolutional neural network was trained via reinforcement learning in a 10 m × 10 m simulated environment to navigate a randomly placed detector-agent to a randomly placed source of varied strength (stopping with perfect knowledge during training). The network agent could move in one of four directions (up, down, left, right) after taking a 1 s count measurement at the current location. During testing, the stopping criteria for this navigational algorithm was based upon a Bayesian likelihood estimation technique of source presence, updating this likelihood after each step, and terminating once the confidence of the source being in a single location exceeded 0.9. A second network was trained and tested with similar architecture as the previous but which contained a fifth action: for self-stopping. The accuracy and speed of localization with set detector and source initializations were compared over 50 trials of MLE-Bayesian approach and 1000 trials of the CNN with self-stopping. The statistical stopping condition yielded a median localization error of ~1.41 m and median localization speed of 12 steps. The machine learning stopping condition yielded a median localization error of 0 m and median localization speed of 17 steps. This work demonstrated two stopping criteria available to a machine learning guided, source localization system.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Line Faults Classification Using Machine Learning on Three Phase Voltages Extracted from Large Dataset of PMU Measurements

An end-to-end supervised learning method is developed to classify transmission line faults in a twoyear field-recorded dataset that includes synchronized measurements of three-phase voltages recorded by 38 Phasor Measurement Units (PMU) sparsely located in in the US Western Grid interconnection. Statistical analysis is performed to extract features from this large dataset to train Support Vector Machine (SVM), Random Forest (RF), and eXtreme Gradient Boosting (XGBoost) classifiers initially. The training further leverages a simulated dataset from a synthetic grid with 12 PMUs to increase the number of faults of types infrequently seen in the field-recorded dataset. Training the classification models with the combined dataset resulted in a classification accuracy of 97.7%. This is a significant improvement over 89.7% to 92.5% accuracy obtained by relying on the field-recorded dataset alone.

47 OTHER INSTRUMENTATION↗

Model orthogonalization and Bayesian forecast mixing via principal component analysis

One can improve predictability in the unknown domain by combining forecasts of imperfect complex computational models using a Bayesian statistical machine learning framework. In many cases, however, the models used in the mixing process are similar. In addition to contaminating the model space, the existence of such similar, or even redundant, models during the multimodeling process can result in misinterpretation of results and deterioration of predictive performance. In this paper we describe a method based on the principal component analysis that eliminates model redundancy. We show that by adding model orthogonalization to the proposed Bayesian model combination framework, one can arrive at better prediction accuracy and reach excellent uncertainty quantification performance.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

BRAKER3: Fully automated genome annotation using RNA-seq and protein evidence with GeneMark-ETP, AUGUSTUS, and TSEBRA

Gene prediction has remained an active area of bioinformatics research for a long time. Still, gene prediction in large eukaryotic genomes presents a challenge that must be addressed by new algorithms. The amount and significance of the evidence available from transcriptomes and proteomes vary across genomes, between genes, and even along a single gene. User-friendly and accurate annotation pipelines that can cope with such data heterogeneity are needed. The previously developed annotation pipelines BRAKER1 and BRAKER2 use RNA-seq or protein data, respectively, but not both. A further significant performance improvement integrating all three data types was made by the recently released GeneMark-ETP. We here present the BRAKER3 pipeline that builds on GeneMark-ETP and AUGUSTUS, and further improves accuracy using the TSEBRA combiner. BRAKER3 annotates protein-coding genes in eukaryotic genomes using both short-read RNA-seq and a large protein database, along with statistical models learned iteratively and specifically for the target genome. We benchmarked the new pipeline on genomes of 11 species under an assumed level of relatedness of the target species proteome to available proteomes. BRAKER3 outperforms BRAKER1 and BRAKER2. The average transcript-level F1-score is increased by about 20 percentage points on average, whereas the difference is most pronounced for species with large and complex genomes. BRAKER3 also outperforms other existing tools, MAKER2, Funannotate, and FINDER. The code of BRAKER3 is available on GitHub and as a ready-to-run Docker container for execution with Docker or Singularity. Overall, BRAKER3 is an accurate, easy-to-use tool for eukaryotic genome annotation.

59 BASIC BIOLOGICAL SCIENCES↗

Bridging the gap between experiments and simulations using machine learning

The physics of inertial confinement fusion is rich and complex. Simulation codes that are used to design experiments are computationally expensive and lack the predictive capability required for extensive parameter exploration in search of a high-performing design for laser direct drive. In this work we use deep learning to build a fast emulator of experiments. To facilitate the development of the deep-learning model, an autoencoder is used to reduce the dimensionality of the input space. Two deep learning models are developed. One model is trained on a vast array of simulation data and is subsequently calibrated to expensive and limited experimental data using a technique known as “transfer learning.” The other model is trained on a statistical model and is subsequently calibrated using experimental data. A comparative study of the two predictive models is carried out. The models potentially reproduce key experimental observables with high accuracy and unprecedented inference times relative to those achieved with simulation codes. These models facilitate rapid exploration of a high dimensional input parameter space.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗