Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Missing data problem”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Suppressing Quantum Circuit Errors Due to System Variability

We present a quantum circuit optimization technique that takes into account the variability in error rates that is inherent across present-day noisy quantum computing platforms. This method can be run after qubit routing or postcompilation and consists of computing isomorphic subgraphs to input circuits and scoring each using heuristic cost functions derived from system calibration data. Using an independent standard algorithmic test suite, we show that it is possible to recover on average nearly 40% of missing fidelity using better qubit selection via efficient to compute cost functions. We demonstrate additional performance gains by considering qubit placement over multiple quantum processors. The overhead from these tools is minimal with respect to other compilation steps, such as qubit routing, as the number of qubits increases. As such, our method can be used to find qubit mappings for problems at the scale of quantum advantage and beyond.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

MLtool: Universal Supervised Machine Learning Tool to Model Tabulated Data

Machine Learning (ML) is a subfield of Artificial Intelligence that gives computers the ability to learn from past data without being explicitly programmed. The predictive capabilities of ML models have already been used to facilitate several scientific breakthroughs. However, the practical application of ML is often limited due to the gaps in technical knowledge of its users. The common issue faced by many scientific researchers is the inability to choose the appropriate ML pipelines that are needed to treat real-world data, which is often sparse and noisy. To solve this problem, we have developed an automated Machine Learning tool (MLtool) that includes a set of ML algorithms and approaches to aid scientific researchers. The current version of MLtool is implemented as an object-oriented Python code that is easily extensible. It includes 44 different regression algorithms used to model data. MLtool helps users select the best model for their data, based on the scoring metrics used. Besides regression algorithms, MLtool also includes a suite of pre- and post-processing techniques such as missing value imputation, categorical variable encoding, input feature normalization, uncertainty quantification, exploratory data analysis (EDA), etc. MLtool was tested on several publicly available multi-dimensional data sets and was found capable of making accurate predictions.

Machine learning↗

MLtool Python Code

Machine Learning (ML) is a subfield of Artificial Intelligence that gives computers the ability to learn from past data without being explicitly programmed. The predictive capabilities of ML models have already been used to facilitate several scientific breakthroughs. However, the practical application of ML is often limited due to the gaps in technical knowledge of its users. The common issue faced by many scientific researchers is the inability to choose the appropriate ML pipelines that are needed to treat real-world data, which is often sparse and noisy. To solve this problem, we have developed an automated Machine Learning tool (MLtool) that includes a set of ML algorithms and approaches to aid scientific researchers. The current version of MLtool is implemented as an object-oriented Python code that is easily extensible. It includes 44 different regression algorithms used to model data. MLtool helps users select the best model for their data, based on the scoring metrics used. Besides regression algorithms, MLtool also includes a suite of pre- and post-processing techniques such as missing value imputation, categorical variable encoding, input feature normalization, uncertainty quantification, exploratory data analysis (EDA), etc. MLtool was tested on several publicly available multi-dimensional data sets and was found capable of making accurate predictions.

Machine Learning↗

ELVES-Dwarf. I. Satellite Systems of Eight Isolated Dwarf Galaxies in the Local Volume

The satellite populations of Milky Way (MW)–mass systems have been extensively studied, significantly advancing our understanding of galaxy formation and dark matter physics. In contrast, the satellites of lower-mass dwarf galaxies remain largely unexplored, despite hierarchical structure formation predicting that dwarf galaxies should host their own satellites. We present the first results of the ELVES-Dwarf survey, which aims to statistically characterize the satellite populations of isolated dwarf galaxies in the Local Volume (4 < D < 10 Mpc). We identify satellite candidates in integrated light using Legacy Surveys data and achieve completeness down to M g ≈ −9 mag. We then confirm the association of satellite candidates with host galaxies using surface-brightness fluctuation distances measured from Hyper Suprime-Cam data. We surveyed eight isolated dwarf galaxies with stellar masses ranging from sub-Small Magellanic Cloud to Large Magellanic Cloud scales $(10^{7.8} < M^{\textrm{host}}_{\star} < 10^{9.5} M_⊙)$, and confirmed six satellites with stellar masses between 10 5.6 and 10 8 M ⊙ . Most confirmed satellites are star-forming, in contrast to the primarily quiescent satellites observed around MW-mass hosts. By comparing observed satellite abundances and stellar mass functions with theoretical predictions, we find no evidence of a “missing satellite problem” in the dwarf galaxy regime.

Li, Jiaxuan 嘉轩李 [Princeton University, NJ (United ↗

Tuning a variational autoencoder for data accountability problem in the Mars Science Laboratory ground data system

The Mars Curiosity rover is frequently sending back engineering and science data that goes through a pipeline of systems before reaching its final destination at the mission operations center making it prone to volume loss and data corruption. A ground data system analysis (GDSA) team is charged with the monitoring of this flow of information and the detection of anomalies in that data in order to request a re-transmission when necessary. This work presents ∆-MADS, a derivative-free optimization method applied for tuning the architecture and hyperparameters of a variational autoencoder trained to detect the data with missing patches in order to assist the GDSA team in their mission.

Lakhmiri, Dounia↗

Nanoscale defect evaluation framework combining real-time transmission electron microscopy and integrated machine learning-particle filter estimation

Observation of dynamic processes by transmission electron microscopy (TEM) is an attractive technique to experimentally analyze materials’ nanoscale phenomena and understand the microstructure-properties relationships in nanoscale. Even if spatial and temporal resolutions of real-time TEM increase significantly, it is still difficult to say that the researchers quantitatively evaluate the dynamic behavior of defects. Images in TEM video are a two-dimensional projection of three-dimensional space phenomena, thus missing information must be existed that makes image’s uniquely accurate interpretation challenging. Therefore, even though they are still a clustering high-dimensional data and can be compressed to two-dimensional, conventional statistical methods for analyzing images may not be powerful enough to track nanoscale behavior by removing various artifacts associated with experiment; and automated and unbiased processing tools for such big-data are becoming mission-critical to discover knowledge about unforeseen behavior. We have developed a method to quantitative image analysis framework to resolve these problems, in which machine learning and particle filter estimation are uniquely combined. The quantitative and automated measurement of the dislocation velocity in an Fe-31Mn-3Al-3Si autunitic steel subjected to the tensile deformation was performed to validate the framework, and an intermittent motion of the dislocations was quantitatively analyzed. The framework is successfully classifying, identifying and tracking nanoscale objects; these are not able to be accurately implemented by the conventional mean-path based analysis.

36 MATERIALS SCIENCE↗

Reference Images from Thin Sections of Lunar Regolith

The specialist literature about the lunar regolith is massive. It is also highly focused on specific topics and effectively impenetrable to most non-geologists. Both characteristics of the literature present substantial hurdles to scientists and engineers interested in the regolith In the author's experience it neither surprising or unusual to find serious misconceptions about lunar-type materials outside of the lunar research community. Education of professionals who are non-geologists but interested in the regolith is impeded by a lack of some basic resources. One asset that has been missing is simply detailed images of the regolith "soil". While a few websites offer imagery of specific features, these are of course selected to illustrate specific features. It is almost impossible for a non-specialist to reason from these what "normal" or "typical" regolith looks like. Further, access to lunar material is highly restricted. And as publications rarely do not provide other than highly focused and narrowly tailored data, there is little potential for workers without personal access to sample to do any work with lunar material. To address both problems the authors have begun to make high resolution optical micrographs of entire thin sections of lunar regolith.

Rickman Doug↗

Floating Block Method for Quantum Monte Carlo Simulations

Quantum Monte Carlo simulations are powerful and versatile tools for the quantum many-body problem. In addition to the usual calculations of energies and eigenstate observables, quantum Monte Carlo simulations can in principle be used to build fast and accurate many-body emulators using eigenvector continuation or design time-dependent Hamiltonians for adiabatic quantum computing. Furthermore, these new applications require something that is missing from the published literature, an efficient quantum Monte Carlo scheme for computing the inner product of ground state eigenvectors corresponding to different Hamiltonians. In this work, we introduce an algorithm called the floating block method, which solves the problem by performing Euclidean time evolution with two different Hamiltonians and interleaving the corresponding time blocks. We use the floating block method and nuclear lattice simulations to build eigenvector continuation emulators for energies of 4 He, 8 Be, 12 C, and 16 O nuclei over a range of local and nonlocal interaction couplings. From the emulator data, we identify the quantum phase transition line from a Bose gas of alpha particles to a nuclear liquid.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Synthetic Foveal Imaging Technology

Synthetic Foveal imaging Technology (SyFT) is an emerging discipline of image capture and image-data processing that offers the prospect of greatly increased capabilities for real-time processing of large, high-resolution images (including mosaic images) for such purposes as automated recognition and tracking of moving objects of interest. SyFT offers a solution to the image-data processing problem arising from the proposed development of gigapixel mosaic focal-plane image-detector assemblies for very wide field-of-view imaging with high resolution for detecting and tracking sparse objects or events within narrow subfields of view. In order to identify and track the objects or events without the means of dynamic adaptation to be afforded by SyFT, it would be necessary to post-process data from an image-data space consisting of terabytes of data. Such post-processing would be time-consuming and, as a consequence, could result in missing significant events that could not be observed at all due to the time evolution of such events or could not be observed at required levels of fidelity without such real-time adaptations as adjusting focal-plane operating conditions or aiming of the focal plane in different directions to track such events. The basic concept of foveal imaging is straightforward: In imitation of a natural eye, a foveal-vision image sensor is designed to offer higher resolution in a small region of interest (ROI) within its field of view. Foveal vision reduces the amount of unwanted information that must be transferred from the image sensor to external image-data-processing circuitry. The aforementioned basic concept is not new in itself: indeed, image sensors based on these concepts have been described in several previous NASA Tech Briefs articles. Active-pixel integrated-circuit image sensors that can be programmed in real time to effect foveal artificial vision on demand are one such example. What is new in SyFT is a synergistic combination of recent advances in foveal imaging, computing, and related fields, along with a generalization of the basic foveal-vision concept to admit a synthetic fovea that is not restricted to one contiguous region of an image.

Hoenk, Michael↗

Hydra: Computer Vision for Online Data Quality Monitoring

Hydra is a system utilizing computer vision for near real-time data quality monitoring. Currently operational across all of Jefferson Lab’s experimental halls, it reduces the workload of shift takers by autonomously monitoring diagnostic plots during experiments. Hydra uses "off-the-shelf" supervised learning technologies and is supported by a comprehensive MySQL database. To simplify access, web apps have been developed to facilitate both labeling and monitoring of Hydra’s inferences. Hydra can connect with the alarm system and incorporates complete historical tracking, enabling it to identify issues that shift takers could miss. When issues are detected, a natural first question is: "Why does Hydra think there is a problem?" To answer, Hydra employs Gradient-weighted Class Activation Maps (GradCAM) to identify regions of the image that are important for the specific classification. This interpretive layer enhances transparency and trustworthiness, which is essential for integration with experiment workflows and operation. The Hydra system, results, and sociological considerations for deployment will be discussed.

Jeske, Torri↗

Aerosol Direct Radiative Effect at the Top of the Atmosphere Over Cloud Free Ocean Derived from Four Years of MODIS Data

A four year record of MODIS spaceborne data provides a new measurement tool to assess the aerosol direct radiative effect at the top of the atmosphere. MODIS derives the aerosol optical thickness and microphysical properties from the scattered sunlight at 0.55-2.1 microns. The monthly MODIS data used here are accumulated measurements across a wide range of view and scattering angles and represent the aerosol s spectrally resolved angular properties. We use these data consistently to compute with estimated accuracy of +/-0.6W/sq m the reflected sunlight by the aerosol over global oceans in cloud free conditions. The MODIS high spatial resolution (0.5 km) allows observation of the aerosol impact between clouds that can be missed by other sensors with larger footprints. We found that over the clear-sky global ocean the aerosol reflected 5.3+/-0.6W/sq m with an average radiative efficiency of 49+/-2W/sq m per unit optical thickness. The seasonal and regional distribution of the aerosol radiative effects are discussed. The analysis adds a new measurement perspective to a climate change problem dominated so far by models.

Remer, L. A.↗

Optical Coatings and Surfaces in Space: MISSE

The space environment presents some unique problems for optics. Components must be designed to survive variations in temperature, exposure to ultraviolet, particle radiation, atomic oxygen and contamination from the immediate environment. To determine the importance of these phenomena, a series of passive exposure experiments have been conducted which included, among others, the Long Duration Exposure Facility (LDEF, 1985- 1990), the Passive Optical Sample Assembly (POSA, 1996- 1997) and most recently, the Materials on the International Space Station Experiment (MISSE, 2001 - 2005). The MISSE program benefited greatly from past experience so that at the conclusion of this 4 year mission, samples which remained intact were in remarkable condition. This study will review data from different aspects of this experiment with emphasis on optical properties and performance.

Stewart, Alan F.↗

Computational Bayesian Methods Applied to Complex Problems in Bio and Astro Statistics

In this dissertation we apply computational Bayesian methods to three distinct problems. In the first chapter, we address the issue of unrealistic covariance matrices used to estimate collision probabilities. We model covariance matrices with a Bayesian Normal-Inverse-Wishart model, which we fit with Gibbs sampling. In the second chapter, we are interested in determining the sample sizes necessary to achieve a particular interval width and establish non-inferiority in the analysis of prevalences using two fallible tests. To this end, we use a third order asymptotic approximation. In the third chapter, we wish to synthesize evidence across multiple domains in measurements taken longitudinally across time, featuring a substantial amount of structurally missing data, and fit the model with Hamiltonian Monte Carlo in a simulation to analyze how estimates of a parameter of interest change across sample sizes.

Elrod, Chris↗

Method for Identifying Probable Archaeological Sites from Remotely Sensed Data

Archaeological sites are being compromised or destroyed at a catastrophic rate in most regions of the world. The best solution to this problem is for archaeologists to find and study these sites before they are compromised or destroyed. One way to facilitate the necessary rapid, wide area surveys needed to find these archaeological sites is through the generation of maps of probable archaeological sites from remotely sensed data. We describe an approach for identifying probable locations of archaeological sites over a wide area based on detecting subtle anomalies in vegetative cover through a statistically based analysis of remotely sensed data from multiple sources. We further developed this approach under a recent NASA ROSES Space Archaeology Program project. Under this project we refined and elaborated this statistical analysis to compensate for potential slight miss-registrations between the remote sensing data sources and the archaeological site location data. We also explored data quantization approaches (required by the statistical analysis approach), and we identified a superior data quantization approached based on a unique image segmentation approach. In our presentation we will summarize our refined approach and demonstrate the effectiveness of the overall approach with test data from Santa Catalina Island off the southern California coast. Finally, we discuss our future plans for further improving our approach.

Tilton, James C.↗

GMI-IPS: Processing & Visualization Software Used in ATom DC-8 Aircraft Studies

NASA's Atmospheric Tomography Mission (ATom) deployed in each of the four seasons during 2016-2018, the DC-8 aircraft in order to establish global-scale datasets intended to improve the representation of chemically reactive gases in global atmospheric chemistry models (ACMs). The Global Modeling Initiative (GMI) executed simulations for each ATom flight using the GMI Chemistry Transport Model (GMI-CTM) to provide species concentrations of chemical gases along the DC-8 flight transects. To solve the problem of translating the GMI-CTM simulation data to the unique spatial resolutions of each ATom flight, the GMI ICARTT Processing Software (GMI-IPS) was developed.The GMI-IPS is written in Python and provides data processing, flight extraction, and visualization support for aircraft research projects using ICARTT format, which is a standard format for airborne instrument data. Additionally, the GMI-IPS interpolates global gridded model data from Hierarchical Data Format (HDF) to ICARTT compatible flight transects. Software classes for instruments and collections provided by the ATom DC-8 aircraft such as MER10, MMS, etc. are derived from a common base class. Other functionality provided by the GMI-IPS are: deriving missing flight entries along a transect, reading ICARTT entries from file, and providing Python data structures for storing flight and model information, and more.The GMI-IPS is GIT source controlled, has approximately 30,000 lines of code, and supports parallelization across data collections. It delivered GMI-CTM data for more than forty distinct DC-8 aircraft flights that took place under ATom. The output ICARTT files adhere to format standard V1.1, and pass the scan utility provided by NASA LaRC Airborne Science Data for Atmospheric Composition. This presentation will include a software and methods overview, and results from ATom, including assessments using the GMI-CTM showing how well observations from ATom flight transects represent a broader region.

Damon, M. R.↗

Learning from Automation Surprises and "Going Sour" Accidents: Progress on Human-Centered Automation

Advances in technology and new levels of automation on commercial jet transports has had many effects. There have been positive effects from both an economic and a safety point of view. The technology changes on the flight deck also have had reverberating effects on many other aspects of the aviation system and different aspects of human performance. Operational experience, research investigations, incidents, and occasionally accidents have shown that new and sometimes surprising problems have arisen as well. What are these problems with cockpit automation, and what should we learn from them? Do they represent over-automation or human error? Or instead perhaps there is a third possibility - they represent coordination breakdowns between operators and the automation? Are the problems just a series of small independent glitches revealed by specific accidents or near misses? Do these glitches represent a few small areas where there are cracks to be patched in what is otherwise a record of outstanding designs and systems? Or do these problems provide us with evidence about deeper factors that we need to address if we are to maintain and improve aviation safety in a changing world? How do the reverberations of technology change on the flight deck provide insight into generic issues about developing human-centered technologies and systems (Winograd and Woods, 1997)? Based on a series of investigations of pilot interaction with cockpit automation (Sarter and Woods, 1992; 1994; 1995; 1997a, 1997 b), supplemented by surveys, operational experience and incident data from other studies (e.g., Degani et al., 1995; Eldredge et al., 1991; Tenney et al., 1995; Wiener, 1989), we too have found that the problems that surround crew interaction with automation are more than a series of individual glitches. These difficulties are symptoms that indicate deeper patterns and phenomena concerning human-machine cooperation and paths towards disaster. In addition, we find the same kinds of patterns behind results from studies of physician interaction with computer-based systems in critical care medicine (e.g., Moll van Charante et al., 1993; Obradovich and Woods, 1996; Cook and Woods, 1996). Many of the results and implications of this kind of research are synthesized and discussed in two comprehensive volumes, Billings (1996) and Woods et al. (1994). This paper summarizes the pattern that has emerged from our research, related research, incident reports, and accident investigations. It uses this new understanding of why problems arise to point to new investment strategies that can help us deal with the perceived "human error" problem, make automation more of a team player, and maintain and improve safety.

Woods, David D.↗

Provider Perspectives: Identification and Follow-up of Infants who Are Deaf or Hard of Hearing

Objective Without timely screening, diagnosis, and intervention, hearing loss can cause significant delays in a child's speech, language, social, and emotional development. In 2019, Texas had nearly twice the average rate of loss to follow-up (LFU) or loss to documentation (LTD; i.e., missing documentation of services received) among infants who did not pass their newborn hearing screening compared to the United States overall (51.1 vs. 27.5%). We aimed to identify factors contributing to LFU/LTD among infants who do not pass their newborn hearing screening in Texas. Study Design Data were collected through semistructured qualitative interviews with 56 providers along the hearing care continuum, including hospital newborn hearing screening program staff, audiologists, primary care physicians, and early intervention (EI) program staff located in three rural and urban public health regions in Texas. Following recording and transcription of the interviews, we used qualitative data analysis software to analyze themes using a conventional content analysis approach. Results Frequently cited barriers included problems with family access to care, difficulty contacting patients, problems with communication between providers and referrals, lack of knowledge among providers and parents, and problems using the online reporting system. Providers in rural areas more often mentioned problems with family access to care and contacting families compared to providers in urban areas. Conclusion These findings provide insight into strategies that public health professionals and health care providers can use to work together to help further increase the number of children identified early who may benefit from EI services. Key Points

Obstetrics & Gynecology↗

Decentralized Low-Rank State Estimation for Power Distribution Systems

This article considers the low-observability state estimation problem in power distribution networks and develops a decentralized state estimation algorithm leveraging the matrix completion methodology. Matrix completion has been shown to be an effective technique in state estimation that exploits the low dimensionality of the power system measurements to recover missing information. This technique can utilize an approximate (linear) load flow model, or it can be used with no physical models in a network where no information about the topology or line admittance is available. The direct application of matrix completion algorithms requires solving a semi-definite programming (SDP) problem, which becomes computationally challenging for large networks. We therefore develop a decentralized algorithm that capitalizes on the popular proximal alternating direction method of multipliers (proximal ADMM). The method allows us to distribute the computation among different areas of the network, leading to a scalable algorithm. By doing all computations at individual control areas and only communicating with neighboring areas, the algorithm eliminates the need for data to be sent to a central processing unit and thus increases efficiency and contributes to the goal of autonomous control of distribution networks. We illustrate the advantages of the proposed algorithm numerically using standard IEEE test cases.

41 EE - Solar Energy Technologies Office (EE-4S)↗