Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Imputation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

49 records · Page 3

Evidence for the Synthesis of ATP by an F0F1 ATP Synthase in Membrane Vesicles from Halorubrum Saccharovorum

Vesicles prepared in a buffer containing ADP, Mg(2+) and Pi synthesized ATP at an initial rate of 2 nmols/min/mg protein after acidification of the bulk medium (pH 8 (right arrow) 4). The intravesicular ATP concentration reached a steady state after about 30 seconds and slowly declined thereafter. ATP synthesis was inhibited by low concentrations of dicyclohexylcarbodiimide and m-chlorophenylhydrazone indicating that synthesis took place in response to the proton gradient. NEM and PCMS, which inhibit vacuolar ATPases and the vacuolar-like ATPases of extreme halophiles, did not affect ATP synthesis, and, in fact, produced higher steady state levels of ATP. This suggested that two ATPase activities were present, one which catalyzed ATP synthesis and one that caused its hydrolysis. Azide, a specific inhibitor of F0F1 ATP Synthases, inhibited halobacterial ATP synthesis. The distribution of acridine orange as imposed by a delta pH demonstrated that azide inhibition was not due to the collapse of the proton gradient due to azide acting as a protonophore. Such an effect was observed, but only at azide concentrations higher than those that inhibited ATP synthesis. These results confirm the earler observations with cells of H. saccharovorum and other extreme halophiles that ATP synthesis is inconsistent with the operation of a vacuolar-like ATPase. Therefore, the observation that a vacuolar-like enzyme is responsible for ATP synthesis (and which serves as the basis for imputing ATP synthesis to the vacuolar-like ATPases of the extreme halophiles, and the Archaea in general) should be taken with some degree of caution.

Faguy, David↗

Making the most of missing values : object clustering with partial data in astronomy

We demonstrate a clustering analysis algorithm, KSC, that a) uses all observed values and b) does not discard the partially observed objects. KSC uses soft constraints defined by the fully observed objects to assist in the grouping of objects with missing values. We present an analysis of objects taken from the Sloan Digital Sky Survey to demonstrate how imputing the values can be misleading and why the KSC approach can produce more appropriate results.

clustering↗

Diurnal Variability of Vertical Structure from a TRMM Passive Microwave "Virtual Radar" Retrieval

Robust description of the diurnal cycle from TRMM observations is complicated by the limitations of Low Earth Orbit (LEO) sampling; from a 'climatological' perspective, sufficient sampling must exist to control for both spatial and seasonal variability, before tackling an additional diurnal component (e.g., with 8 additional 3-hourly or 24 1-hourly bins). For documentation of vertical structure, the narrow sample swath of the TRMM Precipitation Radar limits the resolution of any of these components. A neural-network based 'virtual radar" retrieval has been trained and internally validated, using multifrequency / multipolarization passive microwave(TM1) brightness temperatures and textures parameters and lightning (LIS) observations, as inputs, and PR volumetric reflectivity as targets (outputs). By training the algorithms (essentially highly multivariate, nonlinear regressions) on a very large sample of high-quality co-located data from the center of the TRMM swath, 3D radar reflectivity and derived parameters (VIL, IWC, Echo Tops, etc.) can be retrieved across the entire TMI swath, good to 8-9% over the dynamic range of parameters. As a step in the retrieval (and as an output of the process), each TMI multifrequency pixel (at 85 GHz resolution) is classified into one of the 25 archetypal radar profile vertical structure "types", previously identified using cluster analysis. The dynamic range of retrieved vertical structure appears to have higher fidelity than the current (Version 6) experimental GPROF hydrometeor vertical structure retrievals. This is attributable to correct representation of the prior probabilities of vertical structure variability in the neural network training data, unlike the GPROF cloud-resolving model training dataset used in the V6 algorithms. The LIS lightning inputs are supplementary inputs, and a separate offline neural network has been trained to impute (predict) LIS lightning from passive-microwave-only data. The virtual radar retrieval is thus, in principle, extensible to Aqua/AMSR-E and NPOESS/CMIS passive microwave instruments. The virtual radar approach yields a threefold increase in effective sampling from the mission, albeit of lower-quality "retrieved" data, reducing the variance of local estimates by one third (or the standard deviation by-0.57). In this talk, the variance reduction is leveraged to more finely resolve global diurnal variability in both space and time (local hour).

Boccippio, Dennis J.↗

Affordable Development and Qualification Strategy for Nuclear Thermal Propulsion

A number of recent assessments have confirmed the results of several earlier studies that Nuclear Thermal Propulsion (NTP) is a leading technology for human exploration of Mars. It is generally acknowledged that NTP provides the best prospects for the transportation of humans to Mars in the 2030's. Its high Isp coupled with the high thrusts achievable, allow reasonable trip times, thereby alleviating concerns about space radiation and "claustrophobia" effects. NASA has embarked on the latest phase of the development of NTP systems, and is adopting an affordable approach in response to the pressure of the times. The affordable strategy is built on maximizing the use of the large NTP technology base developed in the 1950's and 60's. The fact that the NTP engines were actually demonstrated to work as planned, is a great risk reduction feature in its development. The strategy utilizes non-nuclear testing to the fullest extent possible, and uses focused nuclear tests for the essential qualification and certification tests. The perceived cost risk of conducting the ground tests is being addressed by considering novel testing approaches. This includes the use of boreholes to contain radioactive effluents, and use of fuel with very high retention capability for fission products. The use of prototype flight tests is being considered as final steps in the development prior to undertaking human flight missions. In addition to the technical issues, plans are being prepared to address the institutional and political issues that need to be considered in this major venture. While the development and deployment of NTP system is not expected to be cheap, the value of the system will be very high, and amortized over the many missions that it enables and enhances, the imputed costs will be very reasonable. Using the approach outlined, NASA and its partners, currently the DOE, and subsequently industry, have a good chance of creating a sustained development program leading to human missions to Mars within the next few decades.

Gerrish, Harold P., Jr.↗

Exploration Analysis of Carbon Dioxide Levels and Ultrasound Measures of the Eye During ISS Missions

Enhanced screening for the Visual Impairment/Intracranial Pressure (VIIP) Syndrome, including in-flight ultrasound, was implemented in 2010 to better characterize the changes in vision observed in some long-duration crewmembers. Suggested possible risk factors for VIIP include cardiovascular changes, diet, anatomical and genetic factors, and environmental conditions. As a potent vasodilator, carbon dioxide (CO (sub 2)), which is chronically elevated on the International Space Station (ISS) relative to typical indoor and outdoor ambient levels on Earth, seems a plausible contributor to VIIP. In an effort to understand the possible associations between CO (sub 2) and VIIP, this study analyzes the relationship between ambient CO (sub 2) levels on ISS and ultrasound measures of the eye obtained from ISS fliers. CO (sub 2) measurements will be pulled directly from Operational Data Reduction Complex for the Lab and Node 3 major constituent analyzers (MCAs) on ISS or from sensors located in the European Columbus module, as available. CO (sub 2) measures between ultrasound sessions will be summarized using standard time series class metrics in MATLAB including time-weighted means and variances. Cumulative CO (sub 2) exposure metrics will also be developed. Regression analyses will be used to quantify the relationships between the CO (sub 2) metrics and specific ultrasound measures. Generalized estimating equations will adjust for the repeated measures within individuals. Multiple imputation techniques will be used to adjust for any possible biases in missing data for either CO (sub 2) or ultrasound measures. These analyses will elucidate the possible relationship between CO (sub 2) and changes in vision and also inform future analysis of inflight VIIP data.

Young, M.↗

Relationship Between Carbon Dioxide Levels and Reported Congestion and Headaches on the International Space Station

Congestion is commonly reported during spaceflight, and most crewmembers have reported using medications for congestion during International Space Station (ISS) missions. Although congestion has been attributed to fluid shifts during spaceflight, fluid status reaches equilibrium during the first week after launch while congestion continues to be reported throughout long duration missions. Congestion complaints have anecdotally been reported in relation to ISS CO2 levels; this evaluation was undertaken to determine whether or not an association exists. METHODS: Reported headaches, congestion symptoms, and CO2 levels were obtained for ISS expeditions 2-31, and time-weighted means and single-point maxima were determined for 24-hour (24hr) and 7-day (7d) periods prior to each weekly private medical conference. Multiple imputation addressed missing data, and logistic regression modeled the relationship between probability of reported event of congestion or headache and CO2 levels, adjusted for possible confounding covariates. The first seven days of spaceflight were not included to control for fluid shifts. Data were evaluated to determine the concentration of CO2 required to maintain the risk of congestion below 1% to allow for direct comparison with a previously published evaluation of CO2 concentrations and headache. RESULTS: This study confirmed a previously identified significant association between CO2 and headache and also found a significant association between CO2 and congestion. For each 1-mm Hg increase in CO2, the odds of a crew member reporting congestion doubled. The average 7-day CO2 would need to be maintained below 1.5 mmHg to keep the risk of congestion below 1%. The predicted probability curves of ISS headache and congestion curves appear parallel when plotted against ppCO2 levels with congestion occurring at approximately 1mmHg lower than a headache would be reported. DISCUSSION: While the cause of congestion is multifactorial, this study showed congestion is associated with CO2 levels on ISS. Data from additional expeditions could be incorporated to further assess this finding. CO2 levels are also associated with reports of headaches on ISS. While it may be expected for astronauts with congestion to also complain of headaches, these two symptoms are commonly mutually exclusive. Furthermore, it is unknown if a temporal CO2 relationship exists between congestion and headache on ISS. CO2 levels were time-weighted for 24hr and 7d, and thus the time course of congestion leading to headache was not assessed; however, congestion could be an early CO2-related symptom when compared to headache. Future studies evaluating the association of CO2-related congestion leading to headache would be difficult due to the relatively stable daily CO2 levels on ISS currently, but a systematic study could be implemented on-orbit if desired.

Cole, Robert↗

Predictive Modeling for Differential Diagnosis and Mortality Risk Assessment

The prevalence of electronic health record (EHR) systems has brought prodigious biomedical informatics opportunity. Automated machine learning methods can effectively utilize such data and have become common tools for healthcare predictive modeling. Researches in medical informatics have explored the potential of deep learning and classical models in emergent care scenarios. In particular, predicting differential diagnoses for admissions have proven useful in decreasing unnecessary lab tests and improving inpatient triage decision-making. Moreover, identification of high-risk patients for in-hospital mortality is vitally important to maximize allocation of medical resources.The Medical Information Mart for Intensive Care (MIMIC-III) database, containing de-identified critical care inpatient was used in our study. This data set captures hospital patient laboratory measurements, pharmacologic prescriptions, diagnostic data and procedure event recordings. When considering adult patients and discounting admissions with ICU length of stay less than 24 hours, there were 37,787 unique admissions and 30,414 total patients. We examined the top 25 most prevalent ICD-9 group-level disease specificities in MIMIC-III using a multi-label classification model. In-hospital mortality was modeled as binary classification with 4,155 (13%) adult patients that expired, of which 3,138 (75.5%) were in the ICU setting. The metrics AUC, F1 score, sensitivity and specificity values calculated for each disease label measured prediction performance.The usage of ICD-9 group codes reduced feature dimension from 14,567 to 942 and greatly improved distribution of patient diagnostic categories. Disease temporal patterns were captured by considering the most frequently sampled 6 vital signs and 13 laboratory values. Missing data were imputed at each time-stamp. Time-series raw hourly average values were converted into 5 summary features (mean, standard deviation, number of observations, min & max values). Patient demographic variables such as age, gender, marital status and ethnicity were also factored into the modeling. Choi et al showed that contextual embedding of medical data, diagnostic and procedural codes alone can predict future diagnoses with sensitivity as high as 0.79. We utilized an embedding technique called word2vec which allowed sparse representations of medical history to be transformed into dense word vectors. The mappings captured contextual information by treating each admission as a sentence and learning the most likely neighboring words in a sliding window fashion. Binary and multi-label classification was achieved via collapse models, which do not consider temporal information, as well as recurrent neural networks with regularization, Softmax output layer activation together with categorical cross-entropy as the loss function.

US Army collaboration↗

Formulative Input into Future NASA Aeronautics Planning

This presentation covers industry input received for future work in NASA Aeronautics over the next 5 years. It is intended to present areas of significant imput and to stimulate further discussion.

future aeronautics planning↗

MLtool: Universal Supervised Machine Learning Tool to Model Tabulated Data

Machine Learning (ML) is a subfield of Artificial Intelligence that gives computers the ability to learn from past data without being explicitly programmed. The predictive capabilities of ML models have already been used to facilitate several scientific breakthroughs. However, the practical application of ML is often limited due to the gaps in technical knowledge of its users. The common issue faced by many scientific researchers is the inability to choose the appropriate ML pipelines that are needed to treat real-world data, which is often sparse and noisy. To solve this problem, we have developed an automated Machine Learning tool (MLtool) that includes a set of ML algorithms and approaches to aid scientific researchers. The current version of MLtool is implemented as an object-oriented Python code that is easily extensible. It includes 44 different regression algorithms used to model data. MLtool helps users select the best model for their data, based on the scoring metrics used. Besides regression algorithms, MLtool also includes a suite of pre- and post-processing techniques such as missing value imputation, categorical variable encoding, input feature normalization, uncertainty quantification, exploratory data analysis (EDA), etc. MLtool was tested on several publicly available multi-dimensional data sets and was found capable of making accurate predictions.

Machine learning↗

Satellite Conjunction Assessment Risk Analysis for "Dilution Region" Events: Issues and Operational Approaches

An important activity within Space Traffic Management is the detection and prevention of possible on-orbit collisions between space objects. The principal parameter for assessing collision likelihood is the probability of collision, which is widely accepted among conjunction assessment practitioners; but it possesses a known deficiency in that it can produce a false sense of safety when the orbital position uncertainties for the conjuncting objects are high. The probability of collision is said to be “diluted” in such a situation and to understate the possible risk; certain approaches have been recommended by researchers to provide (largely conservative) risk estimates and remediation methodologies in these cases. The present analysis explores two of the main proposals for quantifying and remediating possible risk in the dilution region and quantifies their operational implications. These implications with regard to imputed additional workload are considerable, especially in anticipating the conjunction event levels expected with the deployment of the USAF Space Fence radar. This effort has been undertaken as part of a larger enterprise that seeks to clarify the philosophical and statistical underpinnings of the conjunction risk assessment process. The analysis presented herein argues that a form of hypothesis testing is implicitly used in conjunction assessment risk analysis, and that there are a number of conceptual and practical reasons for constructing the associated null hypothesis to counsel against a satellite conjunction remediation action. In short, it is concluded that, for the purposes of determining whether a conjunction remediation action should be pursued, dilution-region probabilities of collision should be treated no differently from those produced under other circumstances.

Probability of collision↗

MLtool Python Code

Machine Learning (ML) is a subfield of Artificial Intelligence that gives computers the ability to learn from past data without being explicitly programmed. The predictive capabilities of ML models have already been used to facilitate several scientific breakthroughs. However, the practical application of ML is often limited due to the gaps in technical knowledge of its users. The common issue faced by many scientific researchers is the inability to choose the appropriate ML pipelines that are needed to treat real-world data, which is often sparse and noisy. To solve this problem, we have developed an automated Machine Learning tool (MLtool) that includes a set of ML algorithms and approaches to aid scientific researchers. The current version of MLtool is implemented as an object-oriented Python code that is easily extensible. It includes 44 different regression algorithms used to model data. MLtool helps users select the best model for their data, based on the scoring metrics used. Besides regression algorithms, MLtool also includes a suite of pre- and post-processing techniques such as missing value imputation, categorical variable encoding, input feature normalization, uncertainty quantification, exploratory data analysis (EDA), etc. MLtool was tested on several publicly available multi-dimensional data sets and was found capable of making accurate predictions.

Machine Learning↗

Bayesian Rules of Thumb: Robust Uncertainty Quantification in Early Project Cost Estimation

Systems engineers often make use of cost Rules ofThumb in order to estimate cost during early phases of projectformulation. These Rules of Thumb typically take the form ofa sequence of percentages over which a total cost is allocatedacross NASA WBS elements. Rules of Thumb can then be usedto extrapolate cost from one or more known WBS elements tothe remaining unknown WBS elements, assisting early projectformulation architecture studies (such as those in JPL’s Team Xand A Team).A number of issues can arise when generating and using costRules of Thumb. For example, many records of project costsconsist of incomplete data. Typical methods of dealing withincomplete cost allocation data include (a) ignoring missionswith incomplete data, or (b) taking averages of the non-zero percentagesacross missions, but both of these methods can result inbiased estimates if the existence of incomplete data correlateswith total mission cost or any particular WBS element. Anothercommon example is cost reported in one or more incorrect WBSelements. This is especially prevalent in smaller missions whereit is more common for engineers to perform tasks that fall underthe purview of multiple WBS elements.Furthermore, a Rule of Thumb estimate is typically reported asa point estimate; there is no reported uncertainty around thepercentages used to generate an allocation. Even in the rarecase in which confidence intervals around mean percentages areprovided, there may be positive or negative correlations betweenWBS elements which can skew estimates.Here we attempt to address these problems by formulatingprobabilistic Rules of Thumb in which a distribution of allocationschemes, rather than a single allocation scheme, is generated.We use a bootstrap imputation method to simultaneouslyaccount for uncertainty in the missing data while using allavailable information contained in the dataset. The imputeddatasets are then input into a multivariate Bayesian modelwhich accounts for correlations between WBS elements andproperly accounts for uncertainty in the final Rule of Thumbpercentages and predictions. We describe the mathematicalmodel and provides snippets of R code utilizing the brms(Bayesian Regression Models using Stan) package. To illustratethis model, we generate a Bayesian Level 2 WBS Cost Rule ofThumb for MIDEX (Medium-Class Explorers) missions withdata extracted from NASA’s CADRe. We then compare thismethod’s performance with the classical Rule of Thumb method.

Hooke, Melissa A↗

Characterizing Spatiotemporal Uncertainty in Interpolated Meteorological Data

Interpolated meteorological data invariably contain errors. These errors have structure in time and space, particularly autocorrelation, which can cause the effects of errors to compound when model outputs are aggregated temporally or spatially. One way to account for this uncertainty is with a probabilistic model from which samples can be drawn that are coherent with respect to underlying spatial and temporal covariance structure. This work describes a probabilistic method for spatial interpolation of point-wise meteorological time series. Observational data from weather stations are generally sparse in space and dense in time (but sometimes missing). The method works by projecting time series onto orthogonal basis vectors and spatially interpolating each resulting component independently. Under suitable assumptions, and data transformations to better satisfy those assumptions, Gaussian process regression provides a complete description of the joint predictive distribution over a Gaussian random field. Spatiotemporally coherent realizations are generated as the sum of conditional (spatial) simulations of each orthogonal (temporal) component. Data-derived and generic orthogonal bases are considered. In addition to spatial interpolation, imputation of missing observational data is examined. The method is applied using near-surface air temperature over the Western United States and validated by comparing theoretical versus actual coverage of predictive distributions and analyzing the degree to which spatial and temporal covariance structure is reproduced. Computational considerations, relating to conditional simulation of random fields, are also addressed.

Conor T Doherty↗