Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data analysis methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

The maximum extent of the filaments and sheets in the cosmic web: an analysis of the SDSS DR17

Filaments and sheets are striking visual patterns in cosmic web. The maximum extent of these large-scale structures are difficult to determine due to their structural variety and complexity. We construct a volume-limited sample of galaxies in a cubic region from the SDSS, divide it into smaller subcubes and shuffle them around. We quantify the average filamentarity and planarity in the 3D galaxy distribution as a function of the density threshold and compare them with those from the shuffled realizations of the original data. The analysis is repeated for different shuffling lengths by varying the size of the subcubes. The average filamentarity and planarity in the shuffled data show a significant reduction when the shuffling scales are smaller than the maximum size of the genuine filaments and sheets. We observe a statistically significant reduction in these statistical measures even at a shuffling scale of $\sim 130 \, {{\, \rm Mpc}}$, indicating that the filaments and sheets in three dimensions can extend up to this length scale. They may extend to somewhat larger length scales that are missed by our analysis due to the limited size of the SDSS data cube. We expect to determine these length scales by applying this method to deeper and larger surveys in future.

79 ASTRONOMY AND ASTROPHYSICS↗

A Hybrid Energy System Workflow for Energy Portfolio Optimization

This manuscript develops a workflow, driven by data analytics algorithms, to support the optimization of the economic performance of an Integrated Energy System. The goal is to determine the optimum mix of capacities from a set of different energy producers (e.g., nuclear, gas, wind and solar). A stochastic-based optimizer is employed, based on Gaussian Process Modeling, which requires numerous samples for its training. Each sample represents a time series describing the demand, load, or other operational and economic profiles for various types of energy producers. These samples are synthetically generated using a reduced order modeling algorithm that reads a limited set of historical data, such as demand and load data from past years. Numerous data analysis methods are employed to construct the reduced order models, including, for example, the Auto Regressive Moving Average, Fourier series decomposition, and the peak detection algorithm. All these algorithms are designed to detrend the data and extract features that can be employed to generate synthetic time histories that preserve the statistical properties of the original limited historical data. The optimization cost function is based on an economic model that assesses the effective cost of energy based on two figures of merit: the specific cash flow stream for each energy producer and the total Net Present Value. An initial guess for the optimal capacities is obtained using the screening curve method. The results of the Gaussian Process model-based optimization are assessed using an exhaustive Monte Carlo search, with the results indicating reasonable optimization results. The workflow has been implemented inside the Idaho National Laboratory’s Risk Analysis and Virtual Environment (RAVEN) framework. The main contribution of this study addresses several challenges in the current optimization methods of the energy portfolios in IES: First, the feasibility of generating the synthetic time series of the periodic peak data; Second, the computational burden of the conventional stochastic optimization of the energy portfolio, associated with the need for repeated executions of system models; Third, the inadequacies of previous studies in terms of the comparisons of the impact of the economic parameters. The proposed workflow can provide a scientifically defendable strategy to support decision-making in the electricity market and to help energy distributors develop a better understanding of the performance of integrated energy systems.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Asymptotic inconsistency of the cumulative algorithm for laser-induced damage probability analysis

The “cumulative algorithm” is a data analysis method that has been proposed to provide an objective, nonparametric determination of laser-induced damage probability as a function of fluence from experimental data that contain both damaged sites and undamaged sites (i.e., 1-on-1 or S-on-1 testing protocols). In this work, the limitations of this approach are explored by considering the asymptotic limit of a large number of test sites. It is shown that the cumulative algorithm does not converge to the true probability distribution and significantly underestimates the damage probability near the damage onset. Here, based on the results of this work, the cumulative algorithm is not recommended for accurate estimation of damage probability.

Computational methods↗

Thermally driven phase transition of halide perovskites revealed by big data-powered in situ electron microscopy

Halide perovskites are promising light-absorbing materials for high-efficiency solar cells, while the crystalline phase of halide perovskites may influence the device’s efficiency and stability. In this work, we investigated the thermally driven phase transition of perovskite (CsPbIxBr3—x), which was confirmed by electron diffraction and high-resolution transmission electron microscopy results. CsPbIxBr3—x transitioned from δ phase to α phase when heated, and the γ phase was obtained when the sample was cooled down. The γ phase was stable as long as it was isolated from humidity and air. A template matching-based data analysis method enabled visualization of the thermally driven phase evolution of perovskite during heating. Here, we also proposed a possible atomic movement in the process of phase transition based on our in situ heating experimental data. The results presented here may improve our understanding of the thermally driven phase transition of perovskite as well as provide a protocol for big-data analysis of in situ experiments.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Modelling stellar activity with Gaussian process regression networks

ABSTRACT Stellar photospheric activity is known to limit the detection and characterization of extrasolar planets. In particular, the study of Earth-like planets around Sun-like stars requires data analysis methods that can accurately model the stellar activity phenomena affecting radial velocity (RV) measurements. Gaussian Process Regression Networks (GPRNs) offer a principled approach to the analysis of simultaneous time series, combining the structural properties of Bayesian neural networks with the non-parametric flexibility of Gaussian Processes. Using HARPS-N solar spectroscopic observations encompassing three years, we demonstrate that this framework is capable of jointly modelling RV data and traditional stellar activity indicators. Although we consider only the simplest GPRN configuration, we are able to describe the behaviour of solar RV data at least as accurately as previously published methods. We confirm the correlation between the RV and stellar activity time series reaches a maximum at separations of a few days, and find evidence of non-stationary behaviour in the time series, associated with an approaching solar activity minimum.

Camacho, J. D. (ORCID:0000000151215560)↗

BeyondPlanck: VII. Bayesian estimation of gain and absolute calibration for cosmic microwave background experiments

We present a Bayesian calibration algorithm for cosmic microwave background (CMB) observations as implemented within the global end-to-end BEYONDPLANCK framework and applied to the Planck Low Frequency Instrument (LFI) data. Following the most recent Planck analysis, we decomposed the full time-dependent gain into a sum of three nearly orthogonal components: one absolute calibration term, common to all detectors, one time-independent term that can vary between detectors, and one time-dependent component that was allowed to vary between one-hour pointing periods. Each term was then sampled conditionally on all other parameters in the global signal model through Gibbs sampling. The absolute calibration is sampled using only the orbital dipole as a reference source, while the two relative gain components were sampled using the full sky signal, including the orbital and Solar CMB dipoles, CMB fluctuations, and foreground contributions. We discuss various aspects of the data that influence gain estimation, including the dipole-polarization quadrupole degeneracy and processing masks. Comparing our solution to previous pipelines, we find good agreement in general, with relative deviations of -0.67% (-0.84%) for 30 GHz, 0.12% (-0.04%) for 44 GHz and -0.03% (-0.64%) for 70 GHz, compared to Planck PR4 and Planck 2018, respectively. We note that the BEYONDPLANCK calibration was performed globally, which results in better inter-frequency consistency than previous estimates. Additionally, WMAP observations were used actively in the BEYONDPLANCK analysis, which both breaks internal degeneracies in the Planck data set and results in an overall better agreement with WMAP. Finally, we used a Wiener filtering approach to smoothing the gain estimates. We show that this method avoids artifacts in the correlated noise maps as a result of oversmoothing the gain solution, which is difficult to avoid with methods like boxcar smoothing, as Wiener filtering by construction maintains a balance between data fidelity and prior knowledge. Although our presentation and algorithm are currently oriented toward LFI processing, the general procedure is fully generalizable to other experiments, as long as the Solar dipole signal is available to be used for calibration.

79 ASTRONOMY AND ASTROPHYSICS↗

Quantitative radiography for determining density fluctuations in HED experiments

We have developed a method to extract density fluctuation measurements from x-ray radiographs of high-energy density (HED) instability growth and turbulence experiments. We use this information to calculate density fluctuation statistics for constraining the performance of turbulent mix models in HED systems. The density calculation combines image filtering, removal of systemic effects such as backlighter variation, calculation of transmission across multiple materials, and use of tracer materials to generate an approximate single-material density field. From the density map, we calculate both average density and a variance-like moment b (density-specific-volume covariance), which we compare to our models. We infer both quantities from a single image, which is significantly more information than the historic single scalar mix width measurements. We also develop a method of analyzing simulation outputs that incorporate both the density fluctuation metric from a turbulence model and the bulk material maps from the hydrodynamic code. This analysis helps address the question of how to initialize the simulations for best comparison to data from systems with large separations of scale in the mixing perturbation initial condition. We find that our data analysis method yields 1D average density and b curves with similar morphology and amplitudes as those from preliminary simulation comparisons.

47 OTHER INSTRUMENTATION↗

Picometer-Precision Atomic Position Tracking through Electron Microscopy

The modern aberration-corrected scanning electron microscopes have successfully achieved direct visualization of atomic columns with sub-angstrom resolution. With this significant progress, advanced image quantification and analysis is still at its early stages. In this work, we present the complete pathway for the metrology of atomic resolution STEM images. This includes: 1) tips for acquiring high-quality STEM images; 2) denoising and drift-correction for enhancing measurement accuracy; 3) obtaining initial atom positions; 4) indexing the atoms based on unit cell vectors; 4) quantifying the atom column positions with either 2D-gaussian single peak fitting or 5) multi-peak fitting routines for slightly overlapping atomic columns; 6) quantification of lattice distortion/strain within the crystal structures or at the defects/interfaces where the lattice periodicity is disrupted, and 7) some common methods to visualize and present the analysis. Furthermore, a simple self-developed free MATLAB app (EASY-STEM) with a graphical user interface (GUI) will be introduced that can help with the analysis of STEM images without the need for writing dedicated analysis code or software. The advanced data analysis methods presented here can be applied for the local quantification of defect relaxations, local structural distortions, local phase transformations, and non-centrosymmetry in a wide range of materials.

47 OTHER INSTRUMENTATION↗

Interlaboratory Reproducibility of Contour Method Data in a High Strength Aluminum Alloy

The contour method for residual stress measurement has seen significant development, but an experimental reproducibility study utilizing physical samples has not been published. A double-blind reproducibly study is reported, having scope beginning with EDM cutting and ending with residual stress calculation. A reinforced I-beam sample geometry is identified for its unique residual stress profile when extracted from residual stress bearing quenched aluminum bar (7050-T74). Contour measurements are prescribed on a midplane of symmetry with dimensions 24.0 mm by 50.0 mm. Fourteen identically prepared samples are fabricated from a single long bar with well characterized and uniform residual stress. Five samples throughout the bar are identified for planning measurements to validate sample uniformity and overall suitability of the residual stress field. The planning measurements employ a range of techniques: contour method, neutron diffraction, and hole-drilling. Eight samples are distributed to an international group of participants to execute their standard measurement practice. A double-blind process is followed to provide anonymity. Results are provided by eight participants: six being self-similar and two being quite different, the latter set aside as outliers. An average residual stress field is established from non-outlying results and the spatial distribution of reproducibility standard deviation is determined. The average stress field ranges from -60 to 70 MPa and the reproducibility standard deviation averages 8.1 MPa on the measurement plane. The average reproducibility standard deviation is about 3 × larger for points within 1.0 mm of plane boundaries (17.6 MPa) than for the remaining points (6.1 MPa). Reproducibility standard deviation (among different labs) for contour method residual stress measurement is found to be very similar to repeatability standard deviation (in a single lab) reported in prior work. The reproducibility observed here, for the entire measurement process, is also similar to that found in a prior reproducibility study limited to contour method data analysis.

36 MATERIALS SCIENCE↗

Scattering Calorimeter FY24 Deliverable Report

A simulation-based method has been developed to prototype new detector designs for nuclear data measurements utilizing neutron scattering. This method uses representative physics inputs for signal and background generation, full detector resolution smearing benchmarked by experimental data, and a neutron beam timing simulation to produce analyzable output like a physical measurement. A test case has been studied using a hybrid time-of-flight calorimeter detector for scattering cross-section measurements on 239 Pu with 1-5 MeV incident monoenergetic neutrons. Data analysis methods have been developed to perform event-level particle reconstruction and reaction channel discrimination. This analysis has been used to estimate the capability of the test detector to perform simultaneous scattering and fission cross section measurements, as well as its ability to provide neutron spectra and particle angular information.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A composite likelihood approach for inference under photometric redshift uncertainty

ABSTRACT Obtaining accurately calibrated redshift distributions of photometric samples is one of the great challenges in photometric surveys like LSST, Euclid, HSC, KiDS, and DES. We present an inference methodology that combines the redshift information from the galaxy photometry with constraints from two-point functions, utilizing cross-correlations with spatially overlapping spectroscopic samples, and illustrate the approach on CosmoDC2 simulations. Our likelihood framework is designed to integrate directly into a typical large-scale structure and weak lensing analysis based on two-point functions. We discuss efficient and accurate inference techniques that allow us to scale the method to the large samples of galaxies to be expected in LSST. We consider statistical challenges like the parametrization of redshift systematics, discuss and evaluate techniques to regularize the sample redshift distributions, and investigate techniques that can help to detect and calibrate sources of systematic error using posterior predictive checks. We evaluate and forecast photometric redshift performance using data from the CosmoDC2 simulations, within which we mimic a DESI-like spectroscopic calibration sample for cross-correlations. Using a combination of spatial cross-correlations and photometry, we show that we can provide calibration of the mean of the sample redshift distribution to an accuracy of at least 0.002(1 + z), consistent with the LSST-Y1 science requirements for weak lensing and large-scale structure probes.

(cosmology:) large-scale structure of Universe↗

Physical Interpretation of Early Battery Life Prediction Models

Early battery life prediction models are most useful for R&D if they help us understand the early changes in battery electrochemical response that correspond with long-term degradation and failure. Linear regression models such as Fused lasso and Partial Least Squares can fit coefficients directly to high-dimensional electrochemical data like capacity-voltage and ΔV–state-of-charge, i.e., Q(V) and ΔV(SOC) curves, learning coefficients that can be physically interpreted. We leverage the ISU-ILCC battery aging data set to learn high-dimensional coefficients for early battery life prediction from traditional slow-rate capacity check data, demonstrating learning on Q(V), d Q· d V −1 , and ΔV(SOC) curves. A thorough study on the dependence of coefficient values on train/test size and data preprocessing methods is made, demonstrating the reliability of high-dimensional regression approaches unless very small amounts of data are used for model training. For this data set, coefficients from Q(V) and d Q· d V −1 models highlight changes in electrode stoichiometry due to lithium loss, while ΔV(SOC) coefficients highlight changes in positive electrode diffusivity due to particle cracking as well as electrode stoichiometry shifts. By directly interpreting the coefficients of a regression model, we make physical insights into battery degradation mechanisms without requiring the assumptions of traditional battery data analysis methods.

25 ENERGY STORAGE↗

Accelerating astronomical and cosmological inference with preconditioned Monte Carlo

ABSTRACT We introduce preconditioned Monte Carlo (PMC), a novel Monte Carlo method for Bayesian inference that facilitates efficient sampling of probability distributions with non-trivial geometry. PMC utilizes a Normalizing Flow (NF) in order to decorrelate the parameters of the distribution and then proceeds by sampling from the preconditioned target distribution using an adaptive Sequential Monte Carlo (SMC) scheme. The results produced by PMC include samples from the posterior distribution and an estimate of the model evidence that can be used for parameter inference and model comparison, respectively. The aforementioned framework has been thoroughly tested in a variety of challenging target distributions achieving state-of-the-art sampling performance. In the cases of primordial feature analysis and gravitational wave inference, PMC is approximately 50 and 25 times faster, respectively, than nested sampling (NS). We found that in higher dimensional applications, the acceleration is even greater. Finally, PMC is directly parallelisable, manifesting linear scaling up to thousands of CPUs.

79 ASTRONOMY AND ASTROPHYSICS↗

BEYONDPLANCK III. Commander3

We describe the computational infrastructure for end-to-end Bayesian cosmic microwave background (CMB) analysis implemented by the BeyondPlanck Collaboration. The code is called Commander3. It provides a statistically consistent framework for global analysis of CMB and microwave observations and may be useful for a wide range of legacy, current, and future experiments. The paper has three main goals. Firstly, we provide a high-level overview of the existing code base, aiming to guide readers who wish to extend and adapt the code according to their own needs or re-implement it from scratch in a different programming language. Secondly, we discuss some critical computational challenges that arise within any global CMB analysis framework, for instance in-memory compression of time-ordered data, fast Fourier transform optimization, and parallelization and load-balancing. Thirdly, we quantify the CPU and RAM requirements for the current BEYONDPLANCK analysis, finding that a total of 1.5 TB of RAM is required for efficient analysis and that the total cost of a full Gibbs sample for LFI is 170 CPU-hrs, including both low-level processing and high-level component separation, which is well within the capabilities of current low-cost computing facilities. The existing code base is made publicly available under a GNU General Public Library (GPL) license.

79 ASTRONOMY AND ASTROPHYSICS↗

Comparing gas composition from fast pyrolysis of live foliage measured in bench-scale and fire-scale experiments

Background: Fire models have used pyrolysis data from oxidising and non-oxidising environments for flaming combustion. In wildland fires pyrolysis, flaming and smouldering combustion typically occur in an oxidising environment (the atmosphere). Aims: Using compositional data analysis methods, determine if the composition of pyrolysis gases measured in non-oxidising and ambient (oxidising) atmospheric conditions were similar. Methods: Permanent gases and tars were measured in a fuel-rich (non-oxidising) environment in a flat flame burner (FFB). Permanent and light hydrocarbon gases were measured for the same fuels heated by a fire flame in ambient atmospheric conditions (oxidising environment). Log-ratio balances of the measured gases common to both environments (CO, CO 2 , CH 4 , H 2 , C 6 H 6 O (phenol), and other gases) were examined by principal components analysis (PCA), canonical discriminant analysis (CDA) and permutational multivariate analysis of variance (PERMANOVA). Key results: Mean composition changed between the non-oxidising and ambient atmosphere samples. PCA showed that flat flame burner (FFB) samples were tightly clustered and distinct from the ambient atmosphere samples. CDA found that the difference between environments was defined by the CO-CO 2 log-ratio balance. PERMANOVA and pairwise comparisons found FFB samples differed from the ambient atmosphere samples which did not differ from each other. Conclusion: Relative composition of these pyrolysis gases differed between the oxidising and non-oxidising environments. This comparison was one of the first comparisons made between bench-scale and field scale pyrolysis measurements using compositional data analysis. Implications: These results indicate the need for more fundamental research on the early time-dependent pyrolysis of vegetation in the presence of oxygen.

54 ENVIRONMENTAL SCIENCES↗

The cosmic web around the Coma cluster from constrained cosmological simulations

Galaxy clusters in the Universe occupy the important position of nodes of the cosmic web. They are connected among them by filaments, elongated structures composed of dark matter, galaxies, and gas. The connection of galaxy clusters to filaments is important, as it is related to the process of matter accretion onto the former. For this reason, investigating the connections to the cosmic web of massive clusters, especially well-known ones for which a lot of information is available, is a hot topic in astrophysics. In a previous work, we performed an analysis of the filament connections of the Coma cluster of galaxies, as detected from the observed galaxy distribution. In this work we resort to a numerical simulation whose initial conditions are constrained to reproduce the local Universe, including the region of the Coma cluster to interpret our observations in an evolutionary context. We detect the filaments connected to the simulated Coma cluster and perform an accurate comparison with the cosmic web configuration we detect in observations. We perform an analysis of the halos’ spatial and velocity distributions close to the filaments in the cluster outskirts. We conclude that, although not significantly larger than the average, the flux of accreting matter on the simulated Coma cluster is significantly more collimated close to the filaments with respect to the general isotropic accretion flux. This paper is the first example of such a result and the first installment in a series of publications which will explore the build-up of the Coma cluster system in connection to the filaments of the cosmic web as a function of redshift.

79 ASTRONOMY AND ASTROPHYSICS↗

Are light curve classification metrics good proxies for SN Ia cosmological constraining power?

Context. When selecting a light curve classifier for use as part of a photometric supernova Ia (SN Ia) cosmological analysis, it is common to make decisions based on metrics of classification performance, such as the contamination within the photometrically classified SN Ia sample, rather than a measure of cosmological constraining power. If the former is an appropriate proxy for the latter, this practice would eliminate the computational expense of a full cosmology forecast in the analysis pipeline design process. Aims. This study tests the assumption that light curve classification metrics are an appropriate proxy for cosmology metrics. Methods. We emulated photometric SN Ia cosmology light curve samples with controlled contamination rates of individual contaminant classes and evaluated each of them under a set of classification metrics. We then derived cosmological parameter constraints from all samples under two common analysis approaches and quantified the impact of contamination by each contaminant class on the resulting cosmological parameter estimates. Results. We observe that cosmology metrics are sensitive to both the contamination rate and the class of the contaminating population, whereas the classification metrics are shown to be insensitive to the latter. Conclusions. Based on these findings, we discourage any exclusive reliance on light curve classification-based metrics for analysis design decisions, which (counterintuitively) include but are not limited to the classifier choice. Instead, we recommend optimising science analysis pipeline design choices using a metric of the information gained about the physical parameters of interest.

79 ASTRONOMY AND ASTROPHYSICS↗

Joint inference of multiplicative and additive systematics in galaxy density fluctuations and clustering measurements

Galaxy clustering measurements are a key probe of the matter density field in the Universe. With the era of precision cosmology upon us, surveys rely on precise measurements of the clustering signal for meaningful cosmological analysis. However, the presence of systematic contaminants can bias the observed galaxy number density, and thereby bias the galaxy two-point statistics. As the statistical uncertainties get smaller, correcting for these systematic contaminants becomes increasingly important for unbiased cosmological analysis. We present and validate a new method for understanding and mitigating both additive and multiplicative systematics in galaxy clustering measurements (two-point function) by joint inference of contaminants in the galaxy overdensity field (one-point function) using a maximum-likelihood estimator (MLE). We test this methodology with Kilo-Degree Survey-like mock galaxy catalogues and synthetic systematic template maps. We estimate the cosmological impact of such mitigation by quantifying uncertainties and possible biases in the inferred relationship between the observed and the true galaxy clustering signal. Our method robustly corrects the clustering signal to the sub-percent level and reduces numerous additive and multiplicative systematics from 1.5σ to less than 0.1σ for the scenarios we tested. In addition, we provide an empirical approach to identifying the functional form (additive, multiplicative, or other) by which specific systematics contaminate the galaxy number density. Even though this approach is tested and geared towards systematics contaminating the galaxy number density, the methods can be extended to systematics mitigation for other two-point correlation measurements.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗