Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data analysis methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

A machine learning approach to galaxy properties: joint redshift–stellar mass probability distributions with Random Forest

We demonstrate that highly accurate joint redshift–stellar mass probability distribution functions (PDFs) can be obtained using the Random Forest (RF) machine learning (ML) algorithm, even with few photometric bands available. As an example, we use the Dark Energy Survey (DES), combined with the COSMOS2015 catalogue for redshifts and stellar masses. We build two ML models: one containing deep photometry in the griz bands, and the second reflecting the photometric scatter present in the main DES survey, with carefully constructed representative training data in each case. We validate our joint PDFs for 10 699 test galaxies by utilizing the copula probability integral transform and the Kendall distribution function, and their univariate counterparts to validate the marginals. Benchmarked against a basic set-up of the template-fitting code bagpipes, our ML-based method outperforms template fitting on all of our predefined performance metrics. In addition to accuracy, the RF is extremely fast, able to compute joint PDFs for a million galaxies in just under 6 min with consumer computer hardware. Such speed enables PDFs to be derived in real time within analysis codes, solving potential storage issues. As part of this work we have developed galpro 1, a highly intuitive and efficient python package to rapidly generate multivariate PDFs on-the-fly. galpro is documented and available for researchers to use in their cosmology and galaxy evolution studies.

79 ASTRONOMY AND ASTROPHYSICS↗

Assessing and benchmarking the fidelity of posterior inference methods for astrophysics data analysis

In this era of large and complex astronomical survey data, interpreting, validating, and comparing inference techniques becomes increasingly difficult. This is particularly critical for emerging inference methods like Simulation-Based Inference (SBI), which offer significant speedup potential and posterior modeling flexibility, especially when deep learning is incorporated. We present a study to assess and compare the performance and uncertainty prediction capability of Bayesian inference algorithms – from traditional MCMC sampling of analytic functions to deep learning-enabled SBI. We focus on testing the capacity of hierarchical inference modeling in those scenarios. Before we extend this study to cosmology, we first use astrophysical simulation data to ensure interpretability. We demonstrate a probabilistic programming implementation of hierarchical and non-hierarchical Bayesian inference using simulations derived from the DeepBench software library, a benchmarking tool developed by our group that generates simple and controllable astrophysical objects from first principles. This study will enable astronomers and physicists to harness the inference potential of these methods with confidence.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Galaxy morphological classification catalogue of the Dark Energy Survey Year 3 data with convolutional neural networks

ABSTRACT We present in this paper one of the largest galaxy morphological classification catalogues to date, including over 20 million galaxies, using the Dark Energy Survey (DES) Year 3 data based on convolutional neural networks (CNNs). Monochromatic i-band DES images with linear, logarithmic, and gradient scales, matched with debiased visual classifications from the Galaxy Zoo 1 (GZ1) catalogue, are used to train our CNN models. With a training set including bright galaxies (16 ≤ i < 18) at low redshift (z < 0.25), we furthermore investigate the limit of the accuracy of our predictions applied to galaxies at fainter magnitude and at higher redshifts. Our final catalogue covers magnitudes 16 ≤ i < 21, and redshifts z < 1.0, and provides predicted probabilities to two galaxy types – ellipticals and spirals (disc galaxies). Our CNN classifications reveal an accuracy of over 99 per cent for bright galaxies when comparing with the GZ1 classifications (i < 18). For fainter galaxies, the visual classification carried out by three of the co-authors shows that the CNN classifier correctly categorizes discy galaxies with rounder and blurred features, which humans often incorrectly visually classify as ellipticals. As a part of the validation, we carry out one of the largest examinations of non-parametric methods, including ∼100 ,000 galaxies with the same coverage of magnitude and redshift as the training set from our catalogue. We find that the Gini coefficient is the best single parameter discriminator between ellipticals and spirals for this data set.

79 ASTRONOMY AND ASTROPHYSICS↗

Preliminary Results from a New Analysis Method for EGRET Data

In order to extend the life of EGRET, the gas in the spark chamber was allowed to deteriorate more than was originally planned for the nominal two year Compton Observatory mission. Gamma ray events are lost because the pattern recognition analysis rules are not optimized for the poorer quality data. By changing the rules used by the data analysts, we can recover a significant fraction of the lost events, allowing improved statistics for detection and study of sources. Preliminary results from the Crab, Geminga, and BL Lacertae indicate the feasibility of this analysis.

Thompson, D. J.↗

Measurement Uncertainty in One-Of-A-Kind Event Data Analysis

A golden standard in science is to repeat an experiment a statistically significant number of times, recording data using the same set of detectors and the same data analysis methodology. In such case experimental error includes both the range of true values generated by repetitions of the experiment, and measurement uncertainty caused by the detector. They are independent. It is a huge and too frequently used simplification, to assume that one can measure multiple repetitions of an identical experiment, resulting in identical true experimental value. Repetitions, as similar is it is experimentally achievable, have unavoidable built-in differences resulting in a range of the true values rather than in a single value. When modern, very sensitive and well calibrated measurement systems are used, this range is not negligible, and sometimes dominates over the measurement uncertainty. Range of true values depends on built-in differences in physics of the experiment. Stochastic physical processes result typically in a broader range of true values than non-stochastic processes do. Measurement uncertainty depends on a measurement method (properties of the detector not of the experiment). Modern measurement methods, including digital ones, frequently make the measurement uncertainty very small. When data from one–of –a kind experiment are analyzed, only the measurement uncertainty is reported. It provides no information about the range of true experimental values, neither about reliability of a reported data point. Reliability of a data point is in general independent from its measurement uncertainty. However, in practice reliable measurement methods frequently have high measurement uncertainty, while low reliability methods are applied to limit measurement uncertainty. Comparison of reliable data with high measurement uncertainty to not so reliable data measured with low uncertainty is discussed – in different scenarios different data analysis methods are applicable. Methods for data analysis from an experiment repeated statistically significant number of times are very well developed. They do not require a detailed expertise in physics of an experiment, nor in the properties of the measurement system used, and meaning of the reported uncertainty is well understood in any scientific community. It all changes when data from one-of-a-kind experiment is analyzed. Analyst’s expertise is required both in the physics of the experiment and in all aspects of the measurement system, all possible malfunctions. Data users must remember that only measurement uncertainty is reported from any one-of-a-kind experiment. Theory with simulations may provide estimation of expected built-in differences in the experiment, and by this of expected range of true values for a given experiment; yet measurement uncertainty can never be used in place of the range of true experimental values.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Advanced statistical methods for improved data analysis of NASA astrophysics missions

The investigators under this grant studied ways to improve the statistical analysis of astronomical data. They looked at existing techniques, the development of new techniques, and the production and distribution of specialized software to the astronomical community. Abstracts of nine papers that were produced are included, as well as brief descriptions of four software packages. The articles that are abstracted discuss analytical and Monte Carlo comparisons of six different linear least squares fits, a (second) paper on linear regression in astronomy, two reviews of public domain software for the astronomer, subsample and half-sample methods for estimating sampling distributions, a nonparametric estimation of survival functions under dependent competing risks, censoring in astronomical data due to nondetections, an astronomy survival analysis computer package called ASURV, and improving the statistical methodology of astronomical data analysis.

Feigelson, Eric D.↗

Comprehensive Data Analysis and Analytic Method Development for PBX9501 Material (FY2019 Annual Report of Aging and Lifetimes Program)

In the polymer bonded explosive (PBX) 9501, a binder consisting of 2.5 wt% nitroplasticizer (NP) and 2.5 wt% Estane R 5703 (Estane) is combined with 94.9 wt% of the high-explosive HMX and 0.1 wt% stabilizer. Because of its flexibility and tensile strength, this binder lowers the sensitivity and improves the manufacturability of PBX 9501. However, like many plasticizers comprised of low molecular weight components, NP has a tendency to diffuse out of the PBX 9501 matrix and can decompose at moderate temperatures into reactive byproducts, such as NO, NO 2 , H 2 O, and HNO x . Through oxidation and hydrolysis, these molecules can further degrade NP and Estane, ultimately degrading the properties of PBX 9501. While gaseous molecules can readily diffuse out of the PBX 9501 charges, the intermediates with lower volatility are more likely trapped in the condensed phase inside the charges and, in a closed system, will co-exist with Estane for an extended period of time creating an ongoing reactive environment. Understanding the rates of formation for these volatiles and intermediates during the NP degradation is therefore a critical prerequisite for understanding the long-term stability of polymeric binders used in munition systems. To work toward this ultimate goal, in the past year, we have conducted comprehensive studies on the physical properties of NP and finished three-year long aging experiment to understand the aging behavior of NP under thermal treatment. Important results are summarized in this annual report. Although NP is widely used in the DOE complex, their physical properties are rather scattered and inconsistent in the open literature. For example, there are at least two widely different values, 14.5°C and -15°C, for the melting point of NP. Although it is known that NP is a eutectic 50:50 mixture of BDNPA and BDNPF, their eutectic phase diagram is not well documented. Furthermore, the effect of temperature on the miscibility between water and eutectic BDNPA/F mixture is rarely reported. To fill these knowledge gaps, a large set of BDNPA/F mixtures with BDNPA concentration ranging from 0 to 100 wt% was analyzed using DSC techniques. In addition to determining the eutectic melt point as -25°C, a phase diagram of the BDNPA/F system was constructed from -30°C to 45°C. With this phase diagram, the phase transition temperatures and composition can be readily found for the BDNPA/F mixtures with various mass ratios.

36 MATERIALS SCIENCE↗

Results of a low power ice protection system test and a new method of imaging data analysis

Tests were conducted on a BF Goodrich De-Icing System's Pneumatic Impulse Ice Protection (PIIP) system in the NASA Lewis Icing Research Tunnel (IRT). Characterization studies were done on shed ice particle size by changing the input pressure and cycling time of the PIIP de-icer. The shed ice particle size was quantified using a newly developed image software package. The tests were conducted on a 1.83 m (6 ft) span, 0.53 m (221 in) chord NACA 0012 airfoil operated at a 4 degree angle of attack. The IRT test conditions were a -6.7 C (20 F) glaze ice, and a -20 C (-4 F) rime ice. The ice shedding events were recorded with a high speed video system. A detailed description of the image processing package and the results generated from this analytical tool are presented.

Shin, Jaiwon↗

Results of a low power ice protection system test and a new method of imaging data analysis

Tests were conducted on a BF Goodrich De-Icing System's Pneumatic Impulse Ice Protection (PIIP) system in the NASA Lewis Icing Research Tunnel (IRT). Characterization studies were done on shed ice particle size by changing the input pressure and cycling time of the PIIP de-icer. The shed ice particle size was quantified using a newly developed image software package. The tests were conducted on a 1.83 m (6 ft) span, 0.53 m (221 in) chord NACA 0012 airfoil operated at a 4 degree angle of attack. The IRT test conditions were a -6.7 C (20 F) glaze ice, and a -20 C (-4 F) rime ice. The ice shedding events were recorded with a high speed video system. A detailed description of the image processing package and the results generated from this analytical tool are presented.

Shin, Jaiwon↗

ATMOS data processing and science analysis methods

The atmospheric trace molecule spectroscopy (ATMOS) instrument, a high-speed Fourier transform spectrometer operating in the middle IR (2.2-16 microns), recorded more than 1500 solar spectra at about 0.0105/cm resolution during its first mission onboard the shuttle Challenger in the spring of 1985. These spectra were acquired during high-sun conditions for studies of the solar atmosphere and during low-sun conditions for studies of the earth's upper atmosphere. This paper describes the steps by which the telemetry data were converted into spectra suitable for analysis, the analysis software and methods developed for the atmospheric and solar studies, and the ATMOS data analysis facility.

Norton, Robert H.↗

Characterizing Reactor Operations from Realistic Simulated Environmental Samples: Combining High-Performance Computing and Data Analytics

Environmental sampling is a common technique employed by inspectors and facility operators in nuclear safeguards, proliferation detection, and process monitoring contexts. Interpreting measurements performed on samples or collections of samples and ensuring the information extracted is accurate and precise is difficult. To date, these analyses have relied on simulated data to enable systematic studies; however, these models are inherently limited by the fidelity of the models and the implicit spatial averaging of isotopic composition or other signatures of interest. To advance this capability, we have refined the spatial discretization and expanded the range of physics in the simulation codes we use to perform reactor simulations and depletion calculations. This allows us to generate data that are more representative of real environmental samples, especially for the length scale of the isotopic composition and associated variation. Accordingly, these new data allow a more realistic assessment of traditional and new data analytic analysis methods. Here we present motivation for developing reactor simulations using high-performance computing methods and resources, impacts of these new simulations on our assessment of data analysis and interpretation methods, and initial results of developing and systematically testing data analytic methods designed to overcome the challenges expected of real-world samples. We also quantify the performance of these analyses using defensible statistical methods.

Dayman, Ken J.↗

Demonstration of Wavelet Techniques in the Spectral Analysis of Bypass Transition Data

A number of wavelet-based techniques for the analysis of experimental data are developed and illustrated. A multiscale analysis based on the Mexican hat wavelet is demonstrated as a tool for acquiring physical and quantitative information not obtainable by standard signal analysis methods. Experimental data for the analysis came from simultaneous hot-wire velocity traces in a bypass transition of the boundary layer on a heated flat plate. A pair of traces (two components of velocity) at one location was excerpted. A number of ensemble and conditional statistics related to dominant time scales for energy and momentum transport were calculated. The analysis revealed a lack of energy-dominant time scales inside turbulent spots but identified transport-dominant scales inside spots that account for the largest part of the Reynolds stress. Momentum transport was much more intermittent than were energetic fluctuations. This work is the first step in a continuing study of the spatial evolution of these scale-related statistics, the goal being to apply the multiscale analysis results to improve the modeling of transitional and turbulent industrial flows.

Lewalle, Jacques↗

A generalized method for the identification of aircraft stability and control derivatives from flight test data.

This paper discusses the application of a generalized identification method for flight test data analysis. The method is based on the maximum likelihood (ML) criterion and includes output error and equation error methods as special cases. Both the linear and nonlinear models with and without process noise are considered. The flight test data from lateral maneuvers of HL-10 and M2/F3 lifting bodies are processed to determine the lateral stability and control derivatives, instrumentation accuracies and biases. A comparison is made between the results of the output error method and the generalized ML method for M2/F3 data containing gusts. It is shown that better fits to time histories are obtained by using the generalized ML method.

Mehra, R. K.↗

Maximum likelihood identification of aircraft stability and control derivatives

Application of a generalized identification method to flight test data analysis. The method is based on the maximum likelihood (ML) criterion and includes output error and equation error methods as special cases. Both the linear and nonlinear models with and without process noise are considered. The flight test data from lateral maneuvers of HL-10 and M2/F3 lifting bodies are processed to determine the lateral stability and control derivatives, instrumentation accuracies, and biases. A comparison is made between the results of the output error method and the ML method for M2/F3 data containing gusts. It is shown that better fits to time histories are obtained by using the ML method. The nonlinear model considered corresponds to the longitudinal equations of the X-22 VTOL aircraft. The data are obtained from a computer simulation and contain both process and measurement noise. The applicability of the ML method to nonlinear models with both process and measurement noise is demonstrated.

Mehra, R. K.↗

A Novel Method for Characterizing Spacesuit Mobility through Metabolic Cost

Spacesuit mobility has historically been defined and characterized by a combination of range of motion and joint torque of the individual anatomical joints when performing isolated motions meant to drive that joint only in a given orthogonal plane. While this has been the standard approach for several decades, there are numerous shortcomings that suit designers and engineers would like to see rectified. First, the lack of a standardized method for collecting both range of motion and joint torque translates to many different test setups, procedures and methods of data analysis. Second, all of these previously used methods for data collection lack some degree of repeatability, even within the same test setup and the same conductor; in addition, attempts at higher fidelity data collection techniques require high overhead and cost with minimal improvement. Lastly, isolated motions in standard anatomical planes are not representative of real‐world tasks that a crewmember would be performing during an EVA, be it microgravity or surface exploration based. To address these shortcomings, options are being explored within the Space Suit and Crew Survival Systems Branch to ascertain the feasibility of an alternative approach to defining mobility - one that is more repeatable, lower overhead, and more tied to functional EVA tasks. This paper serves to document the first attempt at such an alternative option - one that looks at the metabolic energy‐cost of a spacesuit. In other words, can we objectively compare the mobility of a spacesuit by evaluating the metabolic cost of that suit to the wearer while performing a battery of functional EVA tasks?

McFarland, Shane↗

A Bayesian approach to strong lens finding in the era of wide-area surveys

ABSTRACT The arrival of the Vera C. Rubin Observatory’s Legacy Survey of Space and Time (LSST), Euclid-Wide and Roman wide-area sensitive surveys will herald a new era in strong lens science in which the number of strong lenses known is expected to rise from $\mathcal {O}(10^3)$ to $\mathcal {O}(10^5)$. However, current lens-finding methods still require time-consuming follow-up visual inspection by strong lens experts to remove false positives which is only set to increase with these surveys. In this work, we demonstrate a range of methods to produce calibrated probabilities to help determine the veracity of any given lens candidate. To do this we use the classifications from citizen science and multiple neural networks for galaxies selected from the Hyper Suprime-Cam survey. Our methodology is not restricted to particular classifier types and could be applied to any strong lens classifier which produces quantitative scores. Using these calibrated probabilities, we generate an ensemble classifier, combining citizen science, and neural network lens finders. We find such an ensemble can provide improved classification over the individual classifiers. We find a false-positive rate of 10−3 can be achieved with a completeness of 46 per cent, compared to 34 per cent for the best individual classifier. Given the large number of galaxy–galaxy strong lenses anticipated in LSST, such improvement would still produce significant numbers of false positives, in which case using calibrated probabilities will be essential for population analysis of large populations of lenses and to help prioritize candidates for follow-up.

79 ASTRONOMY AND ASTROPHYSICS↗

The IPAC Image Subtraction and Discovery Pipeline for the Intermediate Palomar Transient Factory

We describe the near real-time transient-source discovery engine for the intermediate Palomar Transient Factory (iPTF), currently in operations at the Infrared Processing and Analysis Center (IPAC), Caltech. We coin this system the IPAC/iPTF Discovery Engine (or IDE). We review the algorithms used for PSF-matching, image subtraction, detection, photometry, and machine-learned (ML) vetting of extracted transient candidates. We also review the performance of our ML classifier. For a limiting signal-to-noise ratio of 4 in relatively unconfused regions, bogus candidates from processing artifacts and imperfect image subtractions outnumber real transients by approximately equal to 10:1. This can be considerably higher for image data with inaccurate astrometric and/or PSF-matching solutions. Despite this occasionally high contamination rate, the ML classifier is able to identify real transients with an efficiency (or completeness) of approximately equal to 97% for a maximum tolerable false-positive rate of 1% when classifying raw candidates. All subtraction-image metrics, source features, ML probability-based real-bogus scores, contextual metadata from other surveys, and possible associations with known Solar System objects are stored in a relational database for retrieval by the various science working groups. We review our efforts in mitigating false-positives and our experience in optimizing the overall system in response to the multitude of science projects underway with iPTF.

methods: analytical – methods: data analysis –↗