Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Statistical accuracy”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Accurate core excitation and ionization energies from a state-specific coupled-cluster singles and doubles approach

Here, we investigate the use of orbital-optimized references in conjunction with single-reference coupled-cluster theory with single and double substitutions (CCSD) for the study of core excitations and ionizations of 18 small organic molecules, without the use of response theory or equation-of-motion (EOM) formalisms. Three schemes are employed to successfully address the convergence difficulties associated with the coupled-cluster equations, and the spin contamination resulting from the use of a spin symmetry-broken reference, in the case of excitations. In order to gauge the inherent potential of the methods studied, an effort is made to provide reasonable basis set limit estimates for the transition energies. Overall, we find that the two best-performing schemes studied here for ΔCCSD are capable of predicting excitation and ionization energies with errors comparable to experimental accuracies. The proposed ΔCCSD schemes reduces statistical errors against experimental excitation energies by more than a factor of two when compared to the frozen-core core–valence separated (FC-CVS) EOM-CCSD approach – a successful variant of EOM-CCSD tailored towards core excitations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Wavelet flow for extragalactic foreground simulations

Extragalactic foregrounds in cosmic microwave background (CMB) observations are both a source of cosmological and astrophysical information and a nuisance to the CMB. Effective field-level modeling that captures their non-Gaussian statistical distributions is increasingly important for optimal information extraction, particularly given the low-noise observations from current and upcoming experiments. Here, we explore the use of Wavelet Flow (WF) models to tackle the novel task of modeling the field-level probability distributions of multi-component CMB secondaries and foregrounds. Specifically, we jointly train correlated CMB lensing convergence (κ) and cosmic infrared background (CIB) maps with a WF model and obtain a network that statistically recovers the input to high accuracy — the trained network generates samples of κ and CIB fields whose average power spectra are within a few percent of the inputs across all scales, and whose Minkowski functionals are similarly accurate compared to the inputs. Leveraging the multiscale architecture of these models, we fine-tune both the model parameters and the priors at each scale independently, optimizing performance across different resolutions. These results demonstrate that WF models can accurately simulate correlated components of CMB secondaries, supporting improved analysis of cosmological data. Our code and trained models can be found on this GitHub repo.

cosmological simulations↗

Exploiting non-linear scales in galaxy–galaxy lensing and galaxy clustering: A forecast for the dark energy survey

ABSTRACT The combination of galaxy–galaxy lensing (GGL) and galaxy clustering is a powerful probe of low-redshift matter clustering, especially if it is extended to the non-linear regime. To this end, we use an N-body and halo occupation distribution (HOD) emulator method to model the redMaGiC sample of colour-selected passive galaxies in the Dark Energy Survey (DES), adding parameters that describe central galaxy incompleteness, galaxy assembly bias, and a scale-independent multiplicative lensing bias Alens. We use this emulator to forecast cosmological constraints attainable from the GGL surface density profile ΔΣ(rp) and the projected galaxy correlation function wp, gg(rp) in the final (Year 6) DES data set over scales $r_p=0.3\!-\!30.0\, h^{-1} \, \mathrm{Mpc}$. For a $3{{\ \rm per\ cent}}$ prior on Alens we forecast precisions of $1.9{{\ \rm per\ cent}}$, $2.0{{\ \rm per\ cent}}$, and $1.9{{\ \rm per\ cent}}$ on Ωm, σ8, and $S_8 \equiv \sigma _8\Omega _m^{0.5}$, marginalized over all halo occupation distribution (HOD) parameters as well as Alens. Adding scales $r_p=0.3\!-\!3.0\, h^{-1} \, \mathrm{Mpc}$ improves the S8 precision by a factor of ∼1.6 relative to a large scale ($3.0\!-\!30.0\, h^{-1} \, \mathrm{Mpc}$) analysis, equivalent to increasing the survey area by a factor of ∼2.6. Sharpening the Alens prior to $1{{\ \rm per\ cent}}$ further improves the S8 precision to $1.1{{\ \rm per\ cent}}$, and it amplifies the gain from including non-linear scales. Our emulator achieves per cent-level accuracy similar to the projected DES statistical uncertainties, demonstrating the feasibility of a fully non-linear analysis. Obtaining precise parameter constraints from multiple galaxy types and from measurements that span linear and non-linear clustering offers many opportunities for internal cross-checks, which can diagnose systematics and demonstrate the robustness of cosmological results.

79 ASTRONOMY AND ASTROPHYSICS↗

Beyond the Hype: An Evaluation of Commercially Available Machine-Learning-Based Malware Detectors

There is a lack of scientific testing of commercially available malware detectors, especially those that boast accurate classification of never-before-seen (i.e., zero-day) files using machine learning (ML). Consequently, efficacy of malware detectors is opaque, inhibiting end users from making informed decisions and researchers from targeting gaps in current detectors. In this paper, we present a scientific evaluation of four prominent commercial malware detection tools to assist an organization with two primary questions: To what extent do ML-based tools accurately classify previously and never-before-seen files? Is purchasing a network-level malware detector worth the cost? To investigate, we tested each tool against 3,536 total files (2,554 or 72% malicious, 982 or 28% benign) of a variety of file types, including hundreds of malicious zero-days, polyglots, and APT-style files, delivered on multiple protocols. We present statistical results on detection time and accuracy, consider complementary analysis (using multiple tools together), and provide two novel applications of the recent cost-benefit evaluation procedure of Iannacone & Bridges. Although the ML-based tools are more effective at detecting zero-day files and executables, the signature-based tool might still be an overall better option. Both network-based tools provide substantial (simulated) savings when paired with either host tool, yet both show poor detection rates on protocols other than HTTP or SMTP. Our results show that all four tools have near-perfect precision but alarmingly low recall, especially on file types other than executables and office files—37% of malware, including all polyglot files, were undetected. Priorities for researchers and takeaways for end users are given. Code for future use of the cost model is provided.

97 MATHEMATICS AND COMPUTING↗

Multi-Fidelity Active Subspaces for Wind Farm Uncertainty Quantification

Wind plants operate in stochastic environments characterized by complex turbulent flow dynamics and high-dimensional random variables. A key step in uncertainty quantification studies is sensitivity analysis and dimension reduction that can facilitate the development of surrogate models to be used for forward and inverse propagation or optimization under uncertainty. Prior work has shown active subspaces are an effective tool for identifying important directions in the space of stochastic inputs; however, they have only been applied to single-fidelity wind plant models. In this study, we investigate the efficacy of a multi-fidelity active subspace method for analyzing the uncertainty in wind plant power output. The multi-fidelity active subspace estimator offers the promise of increased accuracy in identifying active subspaces as compared to a single-fidelity estimator for the same computational cost, or a reduction in cost for the same accuracy. This makes the study of uncertainty in larger wind plants and with higher fidelity physics tractable. The multi-fidelity active subspace method is applied to gridded and existing wind plant layouts with single and multiple inflow conditions and its performance for surrogate modeling and uncertainty propagation is compared against a single-fidelity active subspace method. This multi-fidelity approach yields substantial computational speedups of 2x - 3.4x across the test cases along with acceptable accuracy in surrogate modeling and computing statistical moments.

active subspace↗

A General Framework to Learn Tertiary Structure for Protein Sequence Characterization

During the past five years, deep-learning algorithms have enabled ground-breaking progress towards the prediction of tertiary structure from a protein sequence. Very recently, we developed SAdLSA, a new computational algorithm for protein sequence comparison via deep-learning of protein structural alignments. SAdLSA shows significant improvement over established sequence alignment methods. In this contribution, we show that SAdLSA provides a general machine-learning framework for structurally characterizing protein sequences. By aligning a protein sequence against itself, SAdLSA generates a fold distogram for the input sequence, including challenging cases whose structural folds were not present in the training set. About 70% of the predicted distograms are statistically significant. Although at present the accuracy of the intra-sequence distogram predicted by SAdLSA self-alignment is not as good as deep-learning algorithms specifically trained for distogram prediction, it is remarkable that the prediction of single protein structures is encoded by an algorithm that learns ensembles of pairwise structural comparisons, without being explicitly trained to recognize individual structural folds. As such, SAdLSA can not only predict protein folds for individual sequences, but also detects subtle, yet significant, structural relationships between multiple protein sequences using the same deep-learning neural network. The former reduces to a special case in this general framework for protein sequence annotation.

59 BASIC BIOLOGICAL SCIENCES↗

Fast response densitometer for measuring liquid density

Densitometer was developed which produces linear voltage proportional to changes in density of flowing liquid hydrogen. Unit has fast response time and good system stability, statistical variation, and thermal equilibrium. System accuracy is 2 percent of total density span. Basic design may be altered to include measurement of other flowing materials.

Source record↗

An experiment to verify that the weak interactions satisfy the strong equivalence principle

The construction of a clock based on the beta decay process is proposed to test for any violations by the weak interaction of the strong equivalence principle bu determining whether the weak interaction coupling constant beta is spatially constant or whether it is a function of gravitational potential (U). The clock can be constructed by simply counting the beta disintegrations of some suitable source. The total number of counts are to be taken a measure of elapsed time. The accuracy of the clock is limited by the statistical fluctuations in the number of counts, N, which is equal to the square root of N. Increasing N gives a corresponding increase in accuracy. A source based on the electron capture process can be used so as to avoid low energy electron discrimination problems. Solid state and gaseous detectors are being considered. While the accuracy of this type of beta decay clock is much less than clocks based on the electromagnetic interaction, there is a corresponding lack of knowledge of the behavior of beta as a function of gravitational potential. No predictions from nonmetric theories as to variations in beta are available as yet, but they may occur at the U/sg C level.

Eby, P. B.↗

Types and Characteristics of Data for Geomagnetic Field Modeling

Given here is material submitted at a symposium convened on Friday, August 23, 1991, at the General Assembly of the International Union of Geodesy and Geophysics (IUGG) held in Vienna, Austria. Models of the geomagnetic field are only as good as the data upon which they are based, and depend upon correct understanding of data characteristics such as accuracy, correlations, systematic errors, and general statistical properties. This symposium was intended to expose and illuminate these data characteristics.

R A Langel↗

Sensitivity of trajectory calculations to the temporal frequency of wind data

A mesoscale primitive equation model is used to create a 36-h simulation of the three-dimensional wind field of an intense maritime extratropical cyclone. The control experiment uses the simulated wind field every 15 min in a trajectory model to calculate back trajectories from various horizontal and vertical positions of interest relative to synoptic features of the storm. The latter trajectories are compared to trajectories that were calculated with the simulated wind data degraded in time to 30 min, 1 h, 3 h, 6h, and 12 h. Various error statistics reveal significant deterioration in trajectory accuracy between trajectories calculated with 1- and 3-h data frequencies. Trajectories calculated with 15-min, 30-min, and 1-h data frequencies yielded similar results, while trajectories calculated with data time frequencies 3 h and greater yielded results with unacceptably large errors.

Doty, Kevin G.↗

Land cover stratification using Landsat Thematic Mapper data in Sahelian and Sudanian woodland and wooded grassland

A standard methodology for thematic mapping of natural vegetation using remotely sensed imagery and digital image processing was modified to account for the spatial and spectral properties of semi-arid landscapes, and tested in study areas in the Sahelian and Sudanian zones, Mali. A principal components transformation of registered wet and dry season Landsat TM images produced a set of synthetic spectral channels differentiating vegetation cover between seasons, and allowed areas with annual grass growth to be distinguished from areas with woody cover. The transformed data were statistically clustered and clusters were assigned to vegetation type and density categories. In a separate step, the images were manually interpreted to differentiate broad soil classes. Four statistics were compared to evaluate the accuracy of the maps based on sample points from air photos. For the relatively detailed categories initially defined, map accuracies were substandard; however, when vegetation density classes were aggregated, overall accuracy was around 90 percent, and class accuracy was greater than 80 percent for most classes. This method is suitable for stratification and inventory of woody biomass at a regional scale in semi-arid woodland and wooded grassland.

Franklin, J.↗

Tropospheric Ozone Near-Nadir-Viewing IR Spectral Sensitivity and Ozone Measurements from NAST-I

Infrared ozone spectra from near nadir observations have provided atmospheric ozone information from the sensor to the Earth's surface. Simulations of the NPOESS Airborne Sounder Testbed-Interferometer (NAST-I) from the NASA ER-2 aircraft (approximately 20 km altitude) with a spectral resolution of 0.25/cm were used for sensitivity analysis. The spectral sensitivity of ozone retrievals to uncertainties in atmospheric temperature and water vapor is assessed in order to understand the relationship between the IR emissions and the atmospheric state. In addition, ozone spectral radiance sensitivity to its ozone layer densities and radiance weighting functions reveals the limit of the ozone profile retrieval accuracy from NAST-I measurements. Statistical retrievals of ozone with temperature and moisture retrievals from NAST-I spectra have been investigated and the preliminary results from NAST-I field campaigns are presented.

Zhou, Daniel K.↗

Deep Interacting Multiple Model Filtering

In this paper, a deep learning-based multiple model estimation framework is presented for the state estimation of hybrid dynamical systems from high dimensional observations such as camera images. A low dimensional vector which represents the measurement of the latent dynamical system and its corresponding variance are learned using a deep encoder neural network. An Interacting Multiple Model (IMM) filter is used to generate the latent state estimates and covariances using multiple dynamical models, which can be learned using backpropagation through time. The state estimates of the dynamical system and the corresponding covariance matrix are generated from the latent state estimates and covariance using a deep decoder neural network. The whole network is trained in an end-to-end manner using a loss function which minimizes the negative log-likelihood of the neural network parameters. Simulation results are presented using a 2D bouncing ball example and estimation error statistics are computed which demonstrates the accuracy and consistency of the estimation.

Ghananeel Rotithor↗

Methods and Results for a Global Precipitation Measurement (GPM) Validation Network Prototype

As one component of a ground validation system to meet requirements for the upcoming Global Precipitation Measurement (GPM) mission, a quasi-operational prototype a system to compare satellite- and ground-based radar measurements has been developed. This prototype, the GPM Validation Network (VN), acquires data from the Precipitation Radar (PR) on the Tropical Rainfall Measuring Mission (TRMM) satellite and from ground radar (GR) networks in the continental U.S. and participating international sites. PR data serve as a surrogate for similar observations from the Dual-frequency Precipitation Radar (DPR) to be present on GPM. Primary goals of the VN prototype are to understand and characterize the variability and bias of precipitation retrievals between the PR and GR in various precipitation regimes at large scales, and to improve precipitation retrieval algorithms for the GPM instruments. The current VN capabilities concentrate on comparisons of the base reflectivity observations between the PR and GR, and include support for rain rate comparisons. The VN algorithm resamples PR and GR reflectivity and other 2-D and 3-D data fields to irregular common volumes defined by the geometric intersection of the instrument observations, and performs statistical comparisons of PR and GR reflectivity and estimated rain rates. Algorithmic biases and uncertainties introduced by traditional data analysis techniques are minimized by not performing interpolation or extrapolation of data to a fixed grid. The core VN dataset consists of WSR-88D GR data and matching PR orbit subset data covering 21 sites in the southeastern U. S., from August, 2006 to the present. On average, about 3.5 overpass events per month for these WSR-88D sites meet VN criteria for significant precipitation, and have matching PR and GR data available. This large statistical sample has allowed the relative calibration accuracy and stability of the individual ground radars, and the quality of the PR reflectivity attenuation correction in convective and stratiform precipitation to be evaluated. We will present results of PR-GR reflectivity and rain rate bias comparisons for each OR site, and for different rain types, for the full data set and as time series. The capabilities of the statistical analysis and vertical cross section tools for display and analysis of individual site overpass event data will be described, and examples of the tools' outputs will be shown.

Morris, Kenneth R.↗

Investigation of Error Patterns in Geographical Databases

The objective of the research conducted in this project is to develop a methodology to investigate the accuracy of Airport Safety Modeling Data (ASMD) using statistical, visualization, and Artificial Neural Network (ANN) techniques. Such a methodology can contribute to answering the following research questions: Over a representative sampling of ASMD databases, can statistical error analysis techniques be accurately learned and replicated by ANN modeling techniques? This representative ASMD sample should include numerous airports and a variety of terrain characterizations. Is it possible to identify and automate the recognition of patterns of error related to geographical features? Do such patterns of error relate to specific geographical features, such as elevation or terrain slope? Is it possible to combine the errors in small regions into an error prediction for a larger region? What are the data density reduction implications of this work? ASMD may be used as the source of terrain data for a synthetic visual system to be used in the cockpit of aircraft when visual reference to ground features is not possible during conditions of marginal weather or reduced visibility. In this research, United States Geologic Survey (USGS) digital elevation model (DEM) data has been selected as the benchmark. Artificial Neural Networks (ANNS) have been used and tested as alternate methods in place of the statistical methods in similar problems. They often perform better in pattern recognition, prediction and classification and categorization problems. Many studies show that when the data is complex and noisy, the accuracy of ANN models is generally higher than those of comparable traditional methods.

Dryer, David↗

FutureTense

Protective vaccines and reliable diagnostics are essential tools for controlling viral diseases. However, the efficacy of these tools can be diminished by mutations in viral genomes. The delay between the emergence of new viral strains and the redesign of vaccines and diagnostics allows for continued viral transmission. Is it possible to address this challenge by computationally predicting viral genome sequence evolution? Can we “future-proof” vaccines and diagnostics by targeting both current and anticipated future sequence variants? While predicting viral evolution is still an unsolved, “grand challenge” problem in biology, the large, and rapidly growing, number of SARS-CoV-2 genome sequences provide an opportunity to quantify the ability of machine learning to predict viral genome sequence evolution. Towards this end, we have developed a simple computational model for predicting viral evolution at the level of individual nucleotides. The key metric for quantifying the per-base, prediction accuracy for viral evolution is the Mann-Whitney U statistic (or, equivalently, the area under the receiver operator curve). Since the Mann-Whitney U statistic is not a differentiable function, existing deep leaning packages (like Pytorch and Keras/TensorFlow) are not useful, as they require that the accuracy metric/objective function be analytically differentiable with respect to the model parameters. To overcome this challenge, we have implemented custom software, “FutureTense”, that can train a machine learning model by maximizing the non-differentiable Mann-Whitney U statistic. This software trains a machine learning model by exploring along the direction of the discrete gradient of the Mann-Whitney U statistic in the model parameter space. Parallel computing and genome sequence-specific optimizations are used to accelerate model training. The resulting machine learning model learns the observed high C->U mutation rates in the SARS-CoV-2 genome (which are potentially induced by host defenses) and provides prediction accuracies that are significantly better than one would expect from random chance. While predicting viral evolution is still quite far from a solved problem, the surprising performance of this simple model gives hope that the accuracy of predicting viral genome evolution can be further increased by more sophisticated approaches.

Gans, Jason↗